© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 10-GEMMASCOPE-2-TRANSCODER-262K
    4. 45553
    Prev
    Next
    INDEX
    Explanations

    **TOP_POSITIVE_LOGITS:** `a`, `an`, `o`, `ingly`, `い`, `ও`, `\`, `eture`, `|)`,`𝗲`**TOP_ACTIVATING_TEXTS:** These texts contain examples like:* `...` **5**. **Run**...* `...` **5**. **`echo ""`**:* `...` **5**. **`"Hello, world!"`**:* `...` **8**. **Liz Bigham*** `...` **7**. **`goto loop`**:* `...` **4**. **Run**:* `...` **2**. **Replace 'u'*** `...` **7**. **Tighten and Adjust**:* `...` **5**. **`EXISTS (SELECT 1 ...)`**:* `...` **4**. **Time-Frequency Domain Features****MAX_ACTIVATING_TOKENS:** `5`, `5`, `5`, `8`, `7`, `4`, `2`, `7`, `5`, `4`**TOKENS_AFTER_MAX_ACTIVATING_TOKEN:** All are `.`**Analysis:**The `MAX_ACTIVATING_TOKENS` are digits (mostly 5, 4, 7, with some 8, 2). Crucially, in *all* cases, the `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` is a period '.'. This indicates the neuron is likely detecting a numbered item or list element, where the number is followed by a period.The `TOP_POSITIVE_LOGITS` show a mix of things that might follow such a number and period:* Articles/prepositions: `a`, `an`, `o`* Word endings: `ingly`* Non-English characters: `い`, `ও`* Special characters/fragments: `\`, `eture`, `|)`,`𝗲`Looking at the `TOP_ACTIVATING_TEXTS`, these numbers followed by periods are often part of numbered lists (e.g., "5. Run", "5. `echo ""`", "8. Liz Bigham", "7. `goto loop`", "5. `EXISTS`", "4. Time-Frequency Domain Features"). The `MAX_ACTIVATING_TOKENS` seem to represent the *number* in these list items. The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` is the separator `.` after the number.The `TOP_POSITIVE_LOGITS` are the *words/tokens that follow the period*. These are diverse, but the core pattern is a numbered list item marker. The non-English characters suggest it's not strictly limited to English text.Therefore, the neuron seems to detect a numbered list item. The explanation should capture this.Possible explanations:* numbered list items* digits followed by period in lists* list item numbers* numbered list markersThe request asks for a "word or phrase, not a complete sentence" and to "just say the pattern itself". "Numbered list markers" fits this. "digits followed by period" is also good but a bit longer. "List item numbers" is concise.Let's check the constraints:- Concise (3 to 20 words): "numbered list markers" is 3 words. "List item numbers" is 3 words.- Could be single word, phrase, pattern: Yes.- About tokens following/preceding: Yes, specifically tokens preceding (digits) and the fixed token following (period).- Avoid listing tokens: Yes.- No "words related to", etc.: Yes.- No "This neuron detects/predicts": Yes.- No "tokens" or "patterns": "markers" is okay, but "list markers" might be too close. "list item numbers" avoids it.- Not capitalize first letter unless proper noun: "numbered list markers" or "list item numbers".- Specific: Yes."List item numbers" seems the most direct and concise. The neuron is activating when it sees the numbers that typically start list items, followed by the period that separates them from the item's content. The associated logits are then predicting the *content* that follows this number+period marker.Let's refine. The neuron itself is likely activating on the *number* that is *part of a list marker*. The logits predict what comes *after* the number and period. So the neuron's *detection* is tied to the number in a list item.Consider the example: `5. Run`- MAX_ACTIVATING_TOKEN: `5`- TOKENS_AFTER_MAX_ACTIVATING_TOKEN: `.`- TOP_POSITIVE_LOGITS predict what comes after `5.`.So, the neuron detects the *number that begins a list item*."list item numbers" is good."numbered list indicators" - also good."number in a list item" - good.The explanation is about *what the neuron detects*. It detects the *number* which is acting as an *indicator* for a list item.Let's re-evaluate the TOP_POSITIVE_LOGITS. They are very diverse. `a`, `an`, `o` suggest the start of text, `い`, `ও` are non-Latin scripts. `eture` could be end of 'feature', 'picture'. `|)` is an emoticon part.This implies the neuron is detecting the *marker for a list item*, and the following logits are predicting *diverse content* that can appear in such list items.So, "list item numbers" or "numbered list markers" or "list item markers" all capture the idea."list item numbers" - 3 words."numbered list markers" - 3 words."list item markers" - 3 words."list item numbers" is perhaps the most direct description of *what is being activated* (the numbers themselves in their context).Let's look at the structure again.MAX_ACTIVATING_TOKENS: `5`TOKENS_AFTER_MAX_ACTIVATING_TOKEN: `.`This combination `5.` is the "marker".The neuron is activated by `5`, `8`, `7`, `4`, `2` *in the context where they are followed by a period*. This context is found in numbered lists.So, the neuron detects *numbers used as list markers*."list item numbers" feels right.Let's think if there's anything simpler or more evocative.The `TOP_POSITIVE_LOGITS`' diversity is a key clue. It means the neuron isn't *predicting a specific word*, but rather indicating *a point where a new item is starting*."list item numbers" - good."numbered list introduction" - also good."list prefixes" - concise.The digits *are* the prefixes for the list items."list item prefixes" seems specific and covers the "number followed by period" pattern.Let's check again.MAX_ACTIVATING_TOKENS: 5, 5, ... (digits)TOKENS_AFTER_MAX_ACTIVATING_TOKEN: . (period)These are numbered list prefixes.TOP_POSITIVE_LOGITS are predictions of what comes *after* the prefix."list item prefixes" - 3 words. Fits constraints.Consider other candidates:"numbered list items" - describes the whole item, not just the marker

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_10_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    রা
    0.86
    ra
    0.86
    than
    0.86
    to
    0.83
     código
    0.80
    so
    0.79
     cuarto
    0.79
    ра
    0.78
     is
    0.77
    taş
    0.77
    POSITIVE LOGITS
    a
    1.13
    an
    0.84
    o
    0.78
    ingly
    0.75
    い
    0.75
    ও
    0.71
     \
    0.68
    eture
    0.68
    |)
    0.67
    𝗲
    0.67
    Activations Density 0.020%

    No Known Activations