INDEX
Explanations
mentions of news or information
references to misinformation and the impact of news on public perception
New Auto-Interp
Negative Logits
Models
-0.66
Physical
-0.65
uania
-0.64
customs
-0.63
Controls
-0.63
"},"
-0.61
Facilities
-0.61
Physical
-0.60
controls
-0.59
Benefits
-0.58
POSITIVE LOGITS
reader
1.07
commenters
1.00
debunk
0.97
readers
0.96
retweet
0.95
Gawker
0.95
rant
0.90
journalistic
0.90
commenter
0.88
redacted
0.87
Activations Density 0.949%