© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-12B
    3. 24-GEMMASCOPE-2-RES-16K
    4. 1621
    Prev
    Next
    INDEX
    Explanations

    This neuron seems to detect or predict terms related to lists or categories, often followed by punctuation or conjunctions indicating continuation. The presence of foreign language terms in `TOP_POSITIVE_LOGITS` suggests a possible link to multilingual contexts or specific grammatical structures used for listing across languages.The `MAX_ACTIVATING_TOKENS` include words like `testing`, `strength`, `illegal`, `interests`, `efficiency`, `geographies`.The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` are predominantly periods ".", with exceptions like "content" and "is".The `TOP_POSITIVE_LOGITS` are words from various languages (Indic scripts, German, Malay) that often function as conjunctions or indicators of lists/categories (e.g., "such as", "and", "etc.", "people who").Combining these:- `MAX_ACTIVATING_TOKENS` are diverse concepts.- Many are followed by a period, suggesting they might be items in a list.- `TOP_POSITIVE_LOGITS` reinforce the idea of listing or categorizing.The pattern is that these neurons might be firing when a specific concept (like `illegal` or `efficiency`) is presented as an item in a list or a point within a larger structure, often marked by suffixes or subsequent phrases that indicate continuation or definition.Looking at the `TOP_ACTIVATING_TEXTS`:- "debugging, and testing."- "resistance, and strength."- "unethical, or illegal."- "appearance, or interests."- "unethical, and illegal content."- "accuracy and efficiency."- "sectors, and geographies"All these examples show the activated token being part of a list of two or more items, often joined by "and" or "or", and usually ending with a period. The `TOP_POSITIVE_LOGITS` are also words that often appear in lists or enumerations.Therefore, the neuron is likely related to enumerations or lists of items/concepts. The foreign words suggest it might be sensitive to how lists are formed or concluded in various languages.Let's refine this. The neuron activates around words that can be items in a list. The tokens following are often separators or end-of-list markers. The TOP_POSITIVE_LOGITS are words that often signify collection or continuation.Possible explanations:- "items in lists" (3 words)- "concepts as list items" (4 words)- "enumerated concepts" (2 words - too short)- "categorized items or concepts" (4 words)- "listing of concepts" (3 words)- "words in bullet points" (4 words) - inferred from structure.The phrase needs to capture *what* it detects/predicts.items in lists

    np_acts-logits-general · gemini-2.5-flash-lite

    The neuron fires on the principal noun at the end of each directory‐style list entry—i.e. the key subject word finishing each bullet or catalog item.

    oai_token-act-pair · o4-miniTriggered by @jyhe0408
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-12b-pt/resid_post/layer_24_width_16k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     или
    0.72
     או
    0.70
     и
    0.69
     και
    0.68
     or
    0.68
     없고
    0.68
     বা
    0.66
    或者
    0.65
     یا
    0.65
     ή
    0.64
    POSITIVE LOGITS
     ஆகியவற்ற
    0.96
     ஆகியவை
    0.93
     ஆகியோர்
    0.76
     সবই
    0.75
     sowie
    0.74
     alike
    0.73
     ஆகிய
    0.73
     zugleich
    0.73
     allemaal
    0.70
     എന്നിവ
    0.70
    Activations Density 0.637%

    No Known Activations