© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 10-GEMMASCOPE-2-TRANSCODER-262K
    4. 225153
    Prev
    Next
    INDEX
    Explanations

    **Explaining Neuron Behavior: Specificity and Pattern Identification**The core task is to distill the neuron's function into a short phrase (3-20 words) by identifying patterns across the provided lists: `TOP_POSITIVE_LOGITS`, `TOP_ACTIVATING_TEXTS`, `MAX_ACTIVATING_TOKENS`, and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`.Let's analyze the data:**1. MAX_ACTIVATING_TOKENS:**- `Coverage`- `Reach`- `Tuesday`- `Beyond`- `fractional`- `Relevance`This list contains words that seem to relate to specific contexts or attributes. `Coverage` and `Reach` appear in descriptions of physical or functional capabilities. `Tuesday` is a day of the week, suggesting temporal contexts. `Beyond` implies scope or extension. `fractional` suggests numerical or divisibility concepts. `Relevance` points to importance or applicability.**2. TOKENS_AFTER_MAX_ACTIVATING_TOKEN:**- `:**` (repeats multiple times)- `Z`- `a`- `,`- `queens`- `part`The consistent `:**` often follows list items or section headers, indicating a structural or formatting pattern. `queens` is a very specific noun, appearing after a `::` following `two`. `part` also appears.**3. TOP_POSITIVE_LOGITS:**- `HTTPS`- `startTime`- `putBoolean`- `biotin`- `penicillin`- `phenoxy`- `Tween`These are technical terms or words related to science/chemistry/computing. `HTTPS`, `startTime`, `putBoolean` are programming/web-related. `biotin`, `penicillin`, `phenoxy`, `Tween` relate to chemistry/biology/materials. The presence of both technical and scientific terms is interesting.**4. TOP_ACTIVATING_TEXTS:**- Mentions "Cost **Coverage**", "Robot **Reach**".- Mentions "Another **Tuesday**, another deluge."- Mentions "The Merge & **Beyond**".- Mentions "no **fractional** part".- Mentions "Health Professional **Relevance**".- Mentions "No **two** queens can be in the same column."This list strongly reinforces the tokens from `MAX_ACTIVATING_TOKENS`. It shows `Coverage` and `Reach` used in descriptive contexts. `Tuesday` marks a specific day. `Beyond` extends scope. `fractional` relates to numbers. `Relevance` is about importance. `two` is a number.**Finding the Pattern:**- The `MAX_ACTIVATING_TOKENS` and `TOP_ACTIVATING_TEXTS` directly show these words appearing in specific descriptive phrases, often defining aspects or properties.- The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` shows that after some of these activating tokens (like `two`), specific items (`queens`) or structural markers (`:**`) follow.- The `TOP_POSITIVE_LOGITS` show a mix of technical and scientific terms, which might suggest the *domain* or *type* of information the neuron is sensitive to when these patterns occur.Consider the examples:- "Cost **Coverage**", "Robot **Reach**": These are defining characteristics.- "Another **Tuesday**": A specific temporal marker.- "The Merge & **Beyond**": Signifies extension or future scope.- "no **fractional** part": Specific numerical property.- "two queens": Specific count + item (from N-Queens problem). The token after `two` is `queens`.- "Health Professional **Relevance**": Importance/applicability.The neuron seems to be active when specific defining terms or markers appear, often followed by more descriptive elements or structural separators. It might be looking for definitions, specific instances, or points within structured information.The presence of `HTTPS`, `startTime`, `putBoolean` in `TOP_POSITIVE_LOGITS` alongside chemical terms is a bit of a wildcard, but the clear patterns in `MAX_ACTIVATING_TOKENS` and `TOP_ACTIVATING_TEXTS` point towards specific, defining terms.Let's look at the structure around `two` and `queens`:- `MAX_ACTIVATING_TOKENS` includes `two`.- `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` includes `queens` right after `two` (in `...**No two queens can be...`).This suggests a pattern of number + specific noun/concept, or defining attribute + description.The explanation needs to be 3-20 words."coverage", "reach", "Tuesday", "beyond", "fractional", "relevance" - these are all rather specific terms that often introduce or define something.The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` like `:**` and `queens` suggest a context where these terms are part of lists, definitions, or specific problems (like N-Queens).The neuron appears to pick up on terms that introduce specific details, attributes, or entities, especially when they are part of a defined scope or list.Let's try to synthesize:- "specific details or attributes" (3 words) - too general.- "defining terms and specific entities" (4 words) - better, captures both aspects.- "introduces specific details and attributes" (5 words) - good.- "identifies specific details, attributes, and terms" (6 words) - also good.Looking at the `TOP_POSITIVE_LOGITS`, there's `HTTPS`, `startTime`, `putBoolean` (programmatic/web), `biotin`, `penicillin` (chemical/biological). This suggests the neuron isn't *just* about abstract definitions, but might be sensitive to these types of terms within programming or scientific contexts.However, the prompt asks to find *patterns in lists*. The strongest patterns are in `MAX_ACTIVATING_TOKENS` and `TOP_ACTIVATING_TEXTS`.Consider the examples where tokens appear:- `Coverage` *in* text defining what's covered.- `Reach` *in* text defining robot capabilities.- `Tuesday` *in* text describing a day.- `Beyond` *in* text discussing future upgrades.- `fractional` *in* text about numbers.- `two` followed by `queens` in N-Queens.These are all about specifying a part, a quantity, a time, a scope, or a characteristic.How about: "specific item, quantity, or attribute"? (5 words)Or: "specific features, quantities, or times" (5 words)Or: "defining specific features or quantities" (4 words)Let's also consider the `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`. The `:**` suggests list items or definitions.The pattern `number word` (like `two queens`) is also present."defining elements, specific counts, or list items" (7 words)"attributes, specific quantities, or list items" (6 words)The pattern seems to be focusing on elements that define a specific quantity, characteristic, or item within a given context, often marking the beginning of a detail or a list item.The `TOP_POSITIVE_LOGITS` are quite mixed, but if we take `HTTPS`, `startTime`, `putBoolean` as programming/web related, and `biotin`, `penicillin` as science related, the neuron might be detecting these *types* of specific

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_10_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    ینه
    0.55
    Го
    0.51
    ГУ
    0.51
    Ç
    0.50
     свар
    0.49
    ્ટ
    0.47
    čk
    0.46
    ตอนนี้
    0.46
    Κ
    0.46
    ayutt
    0.46
    POSITIVE LOGITS
     HTTPS
    0.48
    startTime
    0.46
    putBoolean
    0.44
    را
    0.44
     biotin
    0.43
     penicillin
    0.43
    phenoxy
    0.43
    Tween
    0.42
    ‌ని
    0.42
    mimpi
    0.41
    Activations Density 0.000%

    No Known Activations