INDEX
Explanations
references to bishops and other religious leaders
New Auto-Interp
Negative Logits
Vice
-0.17
vice
-0.16
cz
-0.15
izza
-0.15
nee
-0.14
ektor
-0.14
Shame
-0.14
ores
-0.14
Dia
-0.14
Gus
-0.14
POSITIVE LOGITS
ric
0.31
rics
0.26
rica
0.18
ricks
0.17
esses
0.17
-elect
0.17
illow
0.17
rick
0.16
ston
0.16
stell
0.16
Activations Density 0.029%