INDEX
Explanations
phrases indicating time or frequency
New Auto-Interp
Negative Logits
ansom
-0.17
enet
-0.16
824
-0.16
achs
-0.15
важа
-0.15
Sanayi
-0.15
alth
-0.14
ent
-0.14
onet
-0.14
WARN
-0.14
POSITIVE LOGITS
completion
0.18
hire
0.17
completion
0.16
éĸĢ
0.15
Completion
0.15
hire
0.15
pitch
0.14
Pitch
0.14
Completion
0.14
룸
0.14
Activations Density 0.103%