© Neuronpedia 2026
Privacy & Terms
Blog
GitHub
Slack
Twitter
Contact
Neuronpedia
Jacobian Lens
NEW
Natural Language
Autoencoders
NEW
Assistant Axis
NEW
Circuit Tracer
UPDATE
Releases
Jump To
Search
Steer
SAE Evals
Exports
interp-engine
NEW
Guides
API
Community
Blog
Privacy & Terms
Contact
Sign In
Home
Models
Llama3.1-8B-IT
llama3.1-8b-it
Meta
Jump to Source/SAE
11-resid-post-aa
Source/SAE
Go
Jump to Feature
11-resid-post-aa
Source/SAE
INDEX
Go
Releases
Finding Misaligned Persona Features in Open-Weight Models
September 2025
Andy Arditi
misaligned-persona
Attention Visualizer
HeadVis
(Luger, Kamath et al.)
Find Head By Metric
Metric & Number of Heads
Induction Score
Prev Token Score
Attention Entropy
Self Attention
Top 8
Click head to select
Layer 2
Head 20
Head 22
Layer 5
Head 8
Head 11
Layer 10
Head 14
Layer 15
Head 28
Head 30
Layer 16
Head 20
Select Head Manually
Layer
0
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
Head Index
0
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
Search Explanations
All
By Release
By Model
By Sources
MODEL
Show Dashboards
Hide Dashboards
Search via Inference
?
Resid Post - 131k
All Layers
SEARCH
Run Example Search
Random
🌮 Food
📰 News
📖 Literary
👯 Personal
🧑💻 Programming
🧑🔬 Technical
🧑🏫 Academic
💼 Business
🧑⚖️ Legal
🧑🏫 Educational
🗼 Cultural
Browse
MODEL
Resid Post - 131k
LAYER
Features in
LLAMA3.1-8B-IT
@
11-resid-post-aa
Hover over a feature on the left to preview its details.
Click a feature to lock it and interact with it.