INDEX
Explanations
No Explanations Found
New Auto-Interp
Negative Logits
iege
-0.71
Crus
-0.68
Bohem
-0.68
Wester
-0.67
IBLE
-0.67
Vide
-0.66
Chall
-0.66
Peel
-0.64
gard
-0.64
Crusader
-0.64
POSITIVE LOGITS
fooled
0.77
screwed
0.69
Zip
0.67
onents
0.65
hooked
0.64
umpy
0.61
precincts
0.61
nces
0.60
pta
0.60
aster
0.60
Activations Density 0.000%
No Known Activations
This feature has no known activations.