INDEX
Explanations
references to the word "blo" or variations of it, potentially indicating a focus on a certain brand or notable term
New Auto-Interp
Negative Logits
bes
-0.17
ounder
-0.16
otherwise
-0.15
lej
-0.15
defaults
-0.15
licable
-0.15
na
-0.14
trá»įng
-0.14
nte
-0.14
nes
-0.14
POSITIVE LOGITS
oming
0.23
blo
0.23
Blo
0.21
oding
0.20
oper
0.18
ober
0.16
Blo
0.16
kovi
0.16
oms
0.15
omba
0.15
Activations Density 0.007%