© Neuronpedia 2026
Privacy & Terms
Blog
GitHub
Slack
Twitter
Contact
Neuronpedia
Jacobian Lens
NEW
Natural Language
Autoencoders
NEW
Assistant Axis
NEW
Circuit Tracer
UPDATE
Releases
Jump To
Search
Steer
SAE Evals
Exports
interp-engine
NEW
Guides
API
Community
Blog
Privacy & Terms
Contact
Sign In
Home
GPT2-Small
Transcoders Residuals
8-TRES-DC
2518
Prev
Next
MODEL
8-tres-dc
INDEX
Go
Explanations
symbols or placeholders indicating omitted or censored content
oai_token-act-pair · gpt-4o-mini
Triggered by @bot
No Scores
New Auto-Interp
AutoInterp Type
claude-4-5-haiku
Generate
Interpretable Score
Intrusion Detection
Top Features by Cosine Similarity
Embeds
Show Plots
Show Explanation
Show Activations
Show Test Field
Show Steer
Show Link
IFrame
<iframe src="https://www.neuronpedia.org/gpt2-small/8-tres-dc/2518?embed=true&embedexplanation=true&embedplots=true&embedsteer=true&embedactivations=true&embedlink=true&embedtest=true" title="Neuronpedia" style="height: 300px; width: 540px;"></iframe>
Link
https://www.neuronpedia.org/gpt2-small/8-tres-dc/2518?embed=true&embedexplanation=true&embedplots=true&embedsteer=true&embedactivations=true&embedlink=true&embedtest=true
Not in Any Lists
Add to List
▼
No Comments
ADD
Negative Logits
iple
-0.83
antid
-0.78
Moody
-0.70
Ness
-0.66
ulkan
-0.66
idi
-0.66
psychology
-0.65
skelet
-0.63
denomin
-0.63
ivia
-0.62
POSITIVE LOGITS
�
1.11
�
1.09
�
1.00
�
0.98
��
0.97
�
0.96
�
0.95
�
0.95
�
0.95
�
0.94
Act
ivations
Density 0.102%
Stacked
Snippet
Full
Show Breaks
Hide Breaks
No Known Activations