© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 9-GEMMASCOPE-2-TRANSCODER-262K
    4. 198694
    Prev
    Next
    INDEX
    Explanations

    **Analyzing Neuron Inputs**I'm currently dissecting the neuron's behavior by examining token lists. My focus is on identifying a pattern within `MAX_ACTIVATING_TOKENS` to distill its influence concisely. This involves observing token frequencies and their relationship to activation.* `**` (appears twice, often signifying emphasis or code)* `do`* `names`* `También` (Spanish for "also")* `EMB` (likely related to embeddings)* `ing` (common English suffix)* `with`* `if`4. **Examine `MAX_ACTIVATING_TOKENS`:** * `**` (appears twice, likely emphasis or markdown) * `do` * `names` * `También` (Spanish for "also") * `EMB` (likely short for embedding) * `ing` (common English suffix) * `with` * `if`5. **Examine `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`:** * `hire` * `select` * `###` (separator) * `you` * `"` * `pueden` (Spanish for "can") * `ED` (suffix) * `ieren` (Spanish suffix e.g. 'quieren' - want) * `the` * `"`6. **Examine `TOP_POSITIVE_LOGITS`:** * `Transportation` * `淈` (seems like a specific character/token) * `cond` (likely abbreviation for condition/conditional) * `empleo` (Spanish for "employment") * `fazla` (Turkish for "puts/makes") * `ోతి` (likely part of a word, e.g. Russian) * `чатки` (Russian for "gloves") * `ුණ` (Armenian character) * `energia` (Spanish/Portuguese for "energy") * `ناد` (Persian/Arabic)7. **Examine `TOP_ACTIVATING_TEXTS`:** * "...**hire a qualified electrician.**..." (following `do`) * "...**select "Encrypt LVM"**. " (following `do`) * "...**Choose a passphrase.** This is the password you'll use to..."**Analyzing Neuron Inputs**I'm currently dissecting the neuron's behavior by examining token lists. My focus is on identifying a pattern within `MAX_ACTIVATING_TOKENS` to distill its influence concisely. This involves observing token frequencies and their relationship to activation.* `**` (appears twice): Often used for emphasis.* `do`: Appears in `MAX_ACTIVATING_TOKENS`. Followed by `hire` and `select` in `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`. This suggests a pattern of "do X" or initiating an action.* `names`: Appears in `MAX_ACTIVATING_TOKENS`. The `TOP_POSITIVE_LOGITS` include words like `Transportation`, `empleo` (employment), `energia` (energy), which are often related to concepts or categories rather than specific named entities, but `names` could also relate to entities if they were present.* `También` (Spanish for "also"): Suggests multilingual or Spanish context.* `EMB`: Likely abbreviation for embedding, suggesting a technical or representational aspect.* `ing`: Common English suffix, part of many verbs.* `with`: Common preposition.* `if`: Appears in `MAX_ACTIVATING_TOKENS`. Followed by `you` in `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`. This suggests conditional statements.* **TOP_POSITIVE_LOGITS**: Contains conceptual words like "Transportation", "empleo", "energia", and some foreign language character sequences. This indicates a broad semantic range, potentially covering different domains or languages. The presence of "empleo" and "energia" might suggest a theme of work/activity/resources.* **TOP_ACTIVATING_TEXTS**: * "...**hire a qualified electrician.**..." - Action-oriented. * "...**select "Encrypt LVM"**. " - Action-oriented, technical. * "...**Choose a passphrase.** This is the password..." - Action-oriented, instructional. * "...shopping for or selecting feminine hygiene products..." - Action/selection, consumer context. * "...assign a score between 0-100. 0" - Instructional, numerical. * "...**DO NOT CHANGE THE FORMAT OF THESE EMBEDS, YOU MUST COPY AND PASTE THEM DIRECTLY.**..." - Instruction, technical. * "...The salmon is infused with the subtle citrus notes of Sunpetal blossoms..." - Descriptive, perhaps related to ingredients or features.**Pattern Identification:**1. **Instructional/Action-oriented Verbs:** Many `TOP_ACTIVATING_TEXTS` contain strong imperative verbs or phrases encouraging action: `hire`, `select`, `Choose`, `DO NOT CHANGE`, `assign a score`. The `MAX_ACTIVATING_TOKENS` `do` and `if` also fit this instructional/conditional theme.2. **Technical/Instructional Context:** Phrases like "qualified electrician", "Encrypt LVM", "passphrase", "COMPATIBILITY", "EMBEDS", "COPY AND PASTE", "score between 0-100" point towards technical documentation, instructions, or analytical tasks.3. **Multilingual Hints:** `También` (Spanish) and `pueden` (Spanish) in `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`, along with foreign characters in `TOP_POSITIVE_LOGITS`, suggest the neuron might engage with or distinguish between different languages.4. **Conceptual Groupings:** `TOP_POSITIVE_LOGITS` like `Transportation`, `empleo`, `energia` suggest it might be sensitive to topics related to resources, work, or abstract concepts that could be part of instructions or descriptions.**Synthesizing the Pattern:**The neuron seems to activate when there are instructions, commands, or conditional statements. It appears to be common in texts that provide guidance, technical steps, or require a decision or action. The presence of Spanish words `También` and `pueden` (and other language indicators in logits) suggests it might also be sensitive to multilingual content or instructions in Spanish. The `EMB` token might relate to "embeddings" themselves, common in technical contexts.Let's try to combine 'instructions', 'actions', 'technical' and 'Spanish'.* "technical instructions and Spanish" (5 words) - Good, but maybe misses the 'actions' aspect strongly.* "instructions, actions, and Spanish" (4 words) - Concise.* "commands, choices, and Spanish commands" (4 words) - A bit redundant.* "technical instructions and Spanish commands" (5 words)? No, it's not just commands.* "instructions for actions" (3 words) - Too general.* "instructional prompts and Spanish" (4 words) - "Prompts" is good.* "commands and Spanish commands" (4 words) - Redundant.* "actionable instructions in Spanish" (4 words) - This captures directives and

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_9_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    .
    0.46
     তাঁকে
    0.44
     DOES
    0.44
    PHONY
    0.43
    Adds
    0.42
    ;
    0.42
    name
    0.41
    :
    0.41
     I
    0.41
     ruined
    0.40
    POSITIVE LOGITS
     Transportation
    0.49
    淈
    0.48
     cond
    0.47
     empleo
    0.46
     fazla
    0.46
    ోతి
    0.46
    чатки
    0.46
    ුණ
    0.46
     energia
    0.46
    ناد
    0.45
    Activations Density 0.000%

    No Known Activations