INDEX
Explanations
references to societal roles and statuses within a community
New Auto-Interp
Negative Logits
cco
-0.19
ãĥĨãĥ«
-0.17
ombok
-0.17
ouz
-0.16
asil
-0.15
SizeMode
-0.15
aria
-0.14
ستر
-0.14
arie
-0.14
fix
-0.14
POSITIVE LOGITS
sturdy
0.19
men
0.18
worth
0.17
invalid
0.16
turbulent
0.15
Scotch
0.15
neutr
0.15
persons
0.15
spoilers
0.15
éĢĨ
0.15
Activations Density 0.253%