INDEX
Explanations
references to pets and their relationships with humans
New Auto-Interp
Negative Logits
Morav
-0.17
Arena
-0.16
andon
-0.16
Greenwood
-0.15
agh
-0.15
Pra
-0.14
æļ
-0.14
inosaur
-0.14
orgh
-0.14
Siemens
-0.14
POSITIVE LOGITS
Shadow
0.20
Buttons
0.19
Brut
0.18
Spot
0.18
Spark
0.17
Shadow
0.17
Spirit
0.16
Beans
0.16
Spark
0.16
Beans
0.16
Activations Density 0.162%