© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 31-GEMMASCOPE-2-RES-262K
    4. 110159
    Prev
    Next
    INDEX
    Explanations

    wild with This neuron seems to be identifying parts of phrases, or perhaps structural elements when specific words like "me", "other", "is", "wild", "at", "point" appear, followed by specific grammatical particles or other contextual words.Looking at `MAX_ACTIVATING_TOKENS`: `me`, `other`, `is`, `wild`, `at`, `point`Looking at `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`: `could`, `bad`, `guid` (from guidelines), `dep` (from depended), `often`, `with`, `an`, `of`The `TOP_POSITIVE_LOGITS` are mostly non-Latin script or symbols. This hints that the neuron might be sensitive to these or their presence alongside certain English sequences.Let's try to find a common theme."me" -> "could""other" -> "bad""is" -> "and""wild" -> "guid" (unlikely direct pairing, but "wild" text example)"at" -> "dep""point" -> "often"The examples in `TOP_ACTIVATING_TEXTS` with these tokens include:- "me" in "asks so much of me, getting that healthy upright without too much stress is challenging enough."- "other" in "one party does not axt consequently. It is better to not hurt each other"- "is" in "without too much stress is challenging enough."- "wild" in "Your hair gets… wild with the regular stuff."- "at" in "at an Italian restaurant"- "point" in "points in the trajectory."The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` like `could`, `bad`, `guid`, `dep`, `often`, `with`, `an`, `of` seem to be common grammatical connectors or continuations of phrases.The `TOP_POSITIVE_LOGITS` are very specific symbols and non-Latin characters. This suggests the neuron might be looking for these symbols or script, potentially even when they are mixed with English text.Considering `TOP_POSITIVE_LOGITS` like `(`, `)`, `’`, `⌀`, `+(` and `MAX_ACTIVATING_TOKENS` like `me`, `other`, `is`, `wild`, `at`, `point` and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` like `could`, `bad`, `guid`, `dep`, `often`, `with`, `an`, `of`.The pattern seems to be related to specific punctuation or non-Latin characters, often appearing before or around words that might indicate relationships, states, or descriptions.Let's re-evaluate the `MAX_ACTIVATING_TOKENS` and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` for simple sequential patterns.`me` `could``other` `bad``is` `and``wild` `guid``at` `dep``point` `often`These seem like common word pairings or continuations.The neuron seems to be detecting parts of phrases or common linguistic structures, potentially influenced by the presence of specific non-Latin characters or symbols.Given the `TOP_POSITIVE_LOGITS`, the neuron might be detecting when certain English phrases or structures are adjacent to or mixed with foreign scripts/symbols or specific punctuation.Let's try to generalize:- "me" + "could"- "other" + "bad"- "is" + "and"- "at" + "dep"- "point" + "often"The `TOP_POSITIVE_LOGITS` could suggest it's linking to things like parenthetical comments `( )`, or specific foreign characters `ร์`, `вся/лся`, `याची`, `⌀`, `ॅ`.Maybe the neuron is detecting specific clauses or sub-phrases, especially those that might be parenthetical or use certain punctuation/symbols.Let's consider the most common words in `MAX_ACTIVATING_TOKENS`: `me`, `other`, `is`, `wild`, `at`, `point`.The most common words in `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`: `could`, `bad`, `guid`, `dep`, `often`.Let's look at the combination:`me` `could` -> "me could"`other` `bad` -> "other bad"`is` `and` -> "is and"`wild` `guid` -> "wild guidelines" (text provided "wild with")`at` `dep` -> "at depended"`point` `often` -> "point often"The `TOP_POSITIVE_LOGITS` are crucial. They are non-alphanumeric or foreign script. This suggests the neuron is looking for English phrases when there's a presence of these non-English/symbolic elements.Possible interpretation: phrases adjacent to specific non-English characters or symbols.Or: specific English phrases like "me could", "other bad", "is and", "at dep", "point often".From the examples:- "me, getting that healthy upright without too much stress is challenging enough." - `me` is present. `is` is present.- "Your hair gets… wild with the regular stuff." - `wild` is present.- "It is better to not hurt each other" - `other` is present.The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` are often verbs or general words.The `TOP_POSITIVE_

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/resid_post/layer_31_width_262k_l0_medium
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    investissement
    0.48
     contacto
    0.47
    kunst
    0.45
     Frankreich
    0.45
     textil
    0.45
     控制
    0.44
     вашу
    0.44
     automobil
    0.44
     edhe
    0.44
     brook
    0.43
    POSITIVE LOGITS
    ร์
    0.46
    (
    0.46
    ўся
    0.46
    )
    0.46
     याची
    0.45
    ’)
    0.45
    ⌀
    0.45
    лся
    0.43
    +(
    0.43
    ॅ
    0.43
    Activations Density 0.001%

    No Known Activations