INDEX
Explanations
references to popular dystopian media and literature
New Auto-Interp
Negative Logits
tingham
-0.16
.Unity
-0.15
EAR
-0.15
rv
-0.15
cobra
-0.15
ORED
-0.14
婦
-0.14
ESCO
-0.14
federation
-0.14
ãĥ¥
-0.14
POSITIVE LOGITS
Percy
0.22
Dash
0.20
Hazel
0.20
Dash
0.18
diver
0.17
enie
0.16
Maze
0.16
_frac
0.16
Greg
0.16
Pierce
0.16
Activations Density 0.034%