© Neuronpedia 2026
Privacy & Terms
Blog
GitHub
Slack
Twitter
Contact
Neuronpedia
Jacobian Lens
NEW
Natural Language
Autoencoders
NEW
Assistant Axis
NEW
Circuit Tracer
UPDATE
Releases
Jump To
Search
Steer
SAE Evals
Exports
interp-engine
NEW
Guides
API
Community
Blog
Privacy & Terms
Contact
Sign In
Home
Gemma-2-27B
10-GEMMASCOPE-RES-131K
72826
Prev
Next
MODEL
10-gemmascope-res-131k
Source/SAE
INDEX
Go
Explanations
A followed by specific noun
np_acts-logits-general · gemini-2.5-flash-lite
No Scores
The neuron activates on isolated uppercase “A” tokens used as answer labels or section headings (e.g. the “A” marking an answer or answer‐score annotation).
oai_token-act-pair · o4-mini
Triggered by @jyhe0408
No Scores
uppercase “A” at the start of a word or as a standalone token, often at the beginning of a line or sentence.
oai_token-act-pair · gpt-5
Triggered by @jyhe0408
No Scores
the single capital letter "A" when it appears as a standalone token or label in technical/programming contexts.
oai_token-act-pair · claude-4-5-sonnet
Triggered by @jyhe0408
No Scores
New Auto-Interp
AutoInterp Type
claude-4-5-haiku
Generate
Interpretable Score
Intrusion Detection
Top Features by Cosine Similarity
Configuration
google/gemma-scope-27b-pt-res/layer_10/width_131k
Prompts (Dashboard)
24,576 prompts, 128 tokens each
Dataset (Dashboard)
monology/pile-uncopyrighted
No Configuration Found
Show All
Embeds
Show Plots
Show Explanation
Show Activations
Show Test Field
Show Steer
Show Link
IFrame
<iframe src="https://www.neuronpedia.org/gemma-2-27b/10-gemmascope-res-131k/72826?embed=true&embedexplanation=true&embedplots=true&embedsteer=true&embedactivations=true&embedlink=true&embedtest=true" title="Neuronpedia" style="height: 300px; width: 540px;"></iframe>
Link
https://www.neuronpedia.org/gemma-2-27b/10-gemmascope-res-131k/72826?embed=true&embedexplanation=true&embedplots=true&embedsteer=true&embedactivations=true&embedlink=true&embedtest=true
Not in Any Lists
Add to List
▼
No Comments
ADD
Negative Logits
diarios
-2.03
刳
-2.03
și
-2.02
黝
-1.95
และ
-1.85
produz
-1.80
şi
-1.79
だけです
-1.77
anc
-1.68
いえば
-1.66
POSITIVE LOGITS
have
1.91
也不會
1.83
穩
1.78
1.73
並不是
1.73
6
1.72
There
1.71
Saltar
1.67
from
1.66
渽
1.66
Act
ivations
Density 0.005%
Stacked
Snippet
Full
Show Breaks
Hide Breaks
No Known Activations