INDEX
Explanations
No Explanations Found
New Auto-Interp
Negative Logits
divest
-0.85
owe
-0.77
affirmative
-0.72
hemy
-0.70
undes
-0.68
ilion
-0.65
Owens
-0.63
Americ
-0.63
assert
-0.63
oulos
-0.62
POSITIVE LOGITS
£ı
0.73
mask
0.72
Mandal
0.70
pubs
0.69
rontal
0.68
aldi
0.68
Sto
0.66
MY
0.65
agar
0.65
eral
0.63
Activations Density 0.000%
No Known Activations
This feature has no known activations.