INDEX
Explanations
No Explanations Found
New Auto-Interp
Negative Logits
}}}
-0.76
uring
-0.74
emaker
-0.70
orno
-0.68
ocene
-0.67
dates
-0.65
+#
-0.64
chars
-0.63
elaide
-0.62
adish
-0.61
POSITIVE LOGITS
kinderg
0.69
south
0.69
Sof
0.68
Front
0.66
brother
0.66
alsh
0.65
pupils
0.63
founded
0.62
dinand
0.62
indal
0.61
Activations Density 0.000%
No Known Activations
This feature has no known activations.