© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-12B-IT
    3. 12-GEMMASCOPE-2-RES-16K
    4. 12928
    Prev
    Next
    INDEX
    Explanations

    The neuron seems to activate when specific, often technical or categorical, terms (like 'style', 'code', 'language', 'discussions', 'techniques') appear, particularly when followed by punctuation or phrases that introduce further explanation, code, or subsequent concepts. The `TOP_POSITIVE_LOGITS` are mostly foreign, suggesting it might also activate on non-standard linguistic elements within these contexts.Let's refine the explanation based on these observations. The `MAX_ACTIVATING_TOKENS` are nouns that often serve as headers or categories. The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` are punctuation or common connectors. The `TOP_ACTIVATING_TEXTS` show these nouns leading into explanations, code definitions, or contextual descriptions. The foreign words in `TOP_POSITIVE_LOGITS` might suggest the neuron is sensitive to *how* these categories are introduced, even if the introduction itself is unusual or in a different language.Considering the rules:- Concise (3-20 words)- Find patterns- Don't list all tokens- Don't start with specific phrases- Don't mention "tokens" or "patterns"- Specific, not genericThe core idea is identifying categorized topics or specific modes of discourse, often followed by an elaboration or instantiation.Let's look at the `MAX_ACTIVATING_TOKENS` again:style, code, language, river, discussions, games, techniques, qualities, Complaint, influenceThese are all potential topics, domains, or descriptors.They are often followed by colons, periods, or "to", which introduce definitions, examples, or continuations.Example: "**style**: ...", "**code** to make ...", "**language**", "**river**! ...", "**discussions**. *", "**games**. It's", "**techniques**, tailored for", "**qualities**. Feel", "Complaint –", "influence, Roman"The foreign words in `TOP_POSITIVE_LOGITS` (like הפ, RIDES, caractér, haka, LOP, нуждa, ‌ها,aimana) are a bit of a wildcard, but they *are* listed as highly positive activations. This could mean the neuron is sensitive to *the introduction of these categories using diverse linguistic forms*, or perhaps it captures a broader signal related to domain-specific language or technical explanations, regardless of the exact word, as long as it's in a list context.Let's try to combine these ideas. The neuron seems to be about identifying a category header, often a noun, followed by context or explanation. The foreign words might indicate a sensitivity to variety or to multilingual contexts where these categories appear.If the neuron is identifying the *presence* of lists or structured data where a category is introduced and then explained, that's a key pattern.The `MAX_ACTIVATING_TOKENS` are the *headers* or *keywords*.The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` are the indicators of *continuation/explanation*.A phrase like "list headers followed by explanation" fits this well.list headers followed by explanation

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-12b-it/resid_post/layer_12_width_16k_l0_medium
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     officially
    0.97
     memories
    0.96
     from
    0.93
     recollections
    0.92
     Scholarships
    0.91
     Quatre
    0.89
     recalled
    0.89
     beneficiaries
    0.89
     petitioners
    0.88
     다섯
    0.88
    POSITIVE LOGITS
    P
    0.91
     הפ
    0.88
    RIDES
    0.83
     caractér
    0.80
     haka
    0.80
    LOP
    0.80
     yaar
    0.80
     нужда
    0.79
    ‌ها
    0.78
     dimana
    0.78
    Activations Density 0.219%

    No Known Activations