INDEX
Explanations
references to privilege
mentions of privilege and related concepts
New Auto-Interp
Negative Logits
Estimated
-0.72
Animated
-0.66
ource
-0.64
atom
-0.63
Plot
-0.61
Reported
-0.60
english
-0.59
estimates
-0.59
Fig
-0.59
atl
-0.58
POSITIVE LOGITS
privilege
4.30
privileges
2.40
privileged
2.06
Priv
2.05
ilege
1.97
priv
1.51
Priv
1.39
entitlement
1.27
pleasure
1.21
ileged
1.19
Activations Density 0.010%