© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 15-GEMMASCOPE-2-TRANSCODER-262K
    4. 194543
    Prev
    Next
    INDEX
    Explanations

    number after "2". (after `body`)3 (after `2`)Let's look for patterns:* The `MAX_ACTIVATING_TOKENS` list contains numbers (`2`, `9`, `9`, `2`), prepositions/conjunctions (`to`, `in`), a letter (`W`), a verb (`do`), and a hyphen (`-`).The lists suggest a neuron that potentially identifies numerical sequences and their corresponding context or descriptive words.Let's break it down:* **MAX_ACTIVATING_TOKENS**: Contains numbers (`2`, `9`, `9`, `2`), common words (`to`, `in`, `do`), a letter (`W`), and a separator (`-`). The presence of numbers and words like `to` and `in` is notable.* **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: Shows what follows the activating tokens. We see `-` after numbers (`2`, `9`, `9`), `make` after `to`, `permanently` after `in`, `BE` after `W`, `while` after `do`, `strong` after `-`, `.` after `body`, and `3` after a `2`.* **TOP_POSITIVE_LOGITS**: Contains words like `as`, `patriarch`, `die`, `culture`, `does`, `actúa`. These seem less directly related to the numerical/sequential pattern.* **TOP_ACTIVATING_TEXTS**: This is the most crucial list for finding context. Let's scan for the tokens from `MAX_ACTIVATING_TOKENS` and what follows them: * "``` **Explanation of Fields:** * **`" - Here, `2` might be associated with a URL part like `_2`. * "` **How to Make Coffee:** A Comprehensive Guide!" - Here, `to` is followed by `Make`. This fits the `to make` pattern. * "* **Small amounts of water ice:** Found in permanently shadowed craters." - Here, `in` is followed by `permanently`. This fits the `in permanently` pattern. * "of these, each with its own strengths and weaknesses: **1. Whole Brain Emulation (WBE) / Mind Uploading:** * **Concept:** This involves meticulously scanning a biological brain (potentially post- " - The `W` from `WBE` is followed by `hole`. This might be less direct, or `W` might be activating for concepts starting with `W`. * "things to do while seeking help, and important considerations." - `do` is followed by `while`. This fits the `do while` pattern. * "the body. This is often used for open fractures, severely comminuted fractures, or when internal fixation isn't suitable." - `body` is followed by `.`. This fits the `body.` pattern. * "ISBN: 978-3-527-34230-6 * **Focus:** A comprehensive textbook covering all aspects of MOF" - Contains multiple numbers and `-`. The `9`s and `2`s are part of ISBNs.The strongest pattern emerging is related to specific word sequences and numerical contexts.* `to make`* `in permanently`* `do while`* `body.`* Numerical identifiers (like `_2` or ISBNs)The neuron seems activated by specific word combinations and also appears sensitive to numerical patterns, particularly appearing before punctuation or specific words. The `TOP_POSITIVE_LOGITS` are more abstract and might represent the *function* of these patterns rather than the direct input.Given the request for a concise phrase (3-20 words) without "tokens" or "patterns", and focusing on what it detects/predicts:The phrase should encompass the idea of compound structures involving specific words or numbers.Possible directions:* Phrases starting with numbers.* Phrases involving specific verb-adverb combinations.* Specific structural elements.Considering the `to make`, `in permanently`, `do while` patterns, it looks for specific functional word pairings.The numbers followed by `-` or `.` suggest numerical context or separators.Let's lean into the idea of specific structured phrases.The neuron appears to detect specific, often functional, word pairings and numerical sequences followed by separators or specific words.* "specific phrases and number contexts" - 5 words* "functional word pairings and numerical context" - 5 words* "identifying common phrases and numerical data" - 5 words* "word pairings and number sequences" - 4 wordsThe `TOP_POSITIVE_LOGITS` like `does`, `culture`, `die`, `patriarch` suggest that these identified structures might be used in contexts related to actions, descriptions, or classifications, but the core *detection* is the structural element.Let's look at the `MAX_ACTIVATING_TOKENS` that have clear `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` matches:* `to` -> `make`* `in` -> `permanently`* `do` -> `while`* `body` -> `.`These are specific word pairings. The numbers also seem to be followed by specific things (`-`, `.`, `3`).The neuron is identifying these specific links.A phrase that captures this would be "specific word and number links".Revised thought: The neuron is detecting specific structures formed by words or numbers."specific word pairings and number structures" - 5 words.Let's re-examine `TOP_POSITIVE_LOGITS`: `as`, `consigo`, `patriarch`, `die`, `culture`, `does`, `actúa`.These might hint at the *purpose* of detecting these structures. For example, `does` suggests identifying actions. `culture` and `patriarch` suggest identifying descriptive terms. `actúa` (Spanish for "acts") is similar to `does`.So, it's not *just* the structure, but what the structure *enables*.However, the instructions emphasize finding patterns in the lists provided directly.The most direct pattern is the specific word pairings and numerical contexts.Let's try to be more descriptive of the *type* of pattern."functional word co-occurrences and numerical identifiers" - 5 words."specific word pairings and number contexts" - 5 words.The neuron seems to be activated by common, often functional, collocations (`to make`, `do while`) and numerical data points followed by specific punctuation or characters.Considering the constraints: 3-20 words, no "tokens" or "patterns", specific."specific word combinations and numerical data" - 5 words."actionable word pairs and numeric data" - 5 words.Let's look at how the `MAX_ACTIVATING_TOKENS` and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` align:`2` -> `-``to` -> `make``in` -> `permanently``do` -> `while``body` -> `.`These are very specific pairings. The `TOP_POSITIVE_LOGITS` are more general terms that might *result* from finding these specific links.The core behavior is detecting these specific links.How about: "specific word and number associations"? - 4 words. This is concise and covers both numbers and words.Let

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_15_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    inians
    0.51
    Nonetheless
    0.48
     мелдеш
    0.47
    (/^\
    0.47
     Nonetheless
    0.46
    usel
    0.45
     AdWords
    0.44
    कांनी
    0.43
    жал
    0.43
     Utilisez
    0.43
    POSITIVE LOGITS
     as
    0.53
     consigo
    0.50
     patriarch
    0.49
     die
    0.47
     culture
    0.47
    },
    0.46
     does
    0.46
    culture
    0.46
    云
    0.45
     actúa
    0.44
    Activations Density 0.000%

    No Known Activations