© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B
    3. 31-GEMMASCOPE-2-RES-262K
    4. 14686
    Prev
    Next
    INDEX
    Explanations

    great challenge, go to, beautiful life, walk*Self-correction*: The instruction states to be concise (3 to 20 words) and explain what the neuron detects or predicts by finding patterns. The pattern identified is that these specific tokens (`great`, `go`, `beautiful`, `walk`) are frequently followed by specific other tokens (`challenge`, `to`, `life`, `boldly`) within the activating texts. The explanation should capture this relationship. The provided example explanation format is a phrase.Let's re-evaluate the prompt's constraints:- "concise explanation (3 to 20 words)"- "captures what the neuron detects or predicts by finding patterns in lists."- "look at the lists and find a pattern that helps you explain the behavior of the neuron."- "just say the pattern itself, and do not start with phrases like 'words related to', 'concepts related to', or 'variations of the word'."- "do not start your explanation with 'This neuron detects/predicts'."- "do not mention 'tokens' or 'patterns' in your explanation."- "do not capitalize the first letter unless it is a proper noun."- "explanation should be specific."The pattern is A followed by B.MAX_ACTIVATING_TOKENS: [great, go, the, beautiful, walk]TOKENS_AFTER_MAX_ACTIVATING_TOKEN: [Remove, challenge, to, dull, life, boldly]Connections:great -> challengego -> tothe -> dullbeautiful -> lifewalk -> boldlyThe phrase could describe this pairing."great challenge, go to, beautiful life, walk boldly" - this is a list of phrases showing the pairing. It is 9 words. It is specific. It avoids forbidden phrases.Let's consider the TOP_POSITIVE_LOGITS list again. It contains non-English characters. This might imply the neuron is multilingual or acts on specific character sets. However, the MAX_ACTIVATING_TOKENS and TOKENS_AFTER_MAX_ACTIVATING_TOKEN provide a clear, actionable pattern for English text, which seems to be the primary focus of explanation generation for these types of tasks.If I need to give what the neuron *detects or predicts*, it detects these specific adjacencies or sequential relationships.Alternative interpretation: What is the *concept* represented?"great challenge" - positive action/goal"go to" - movement/destination"beautiful life" - positive description of existence"walk boldly" - confident actionThese might be positive phrases or descriptions of action/state.However, the very specific pairing across the lists is the most obvious pattern.Let's list the pairings clearly:great ... challengego ... tothe ... dullbeautiful ... lifewalk ... boldlyThe phrase "great challenge, go to, beautiful life, walk boldly" exemplifies this.It's a bit long if interpreted as a list of *what it predicts independently*.The prompt says "find a pattern that helps you explain the behavior". The pattern is the co-occurrence or sequential relationship.What if the explanation refers to the *kind* of pairings?"collocations after positive words" - too abstract, uses forbidden words."sequential adverb/verb" - too technical.The most direct way to show the discovered pattern, adhering to the rules, is to list the observed sequences. The user provided `MAX_ACTIVATING_TOKENS` and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`. The pattern is `MAX_ACTIVATING_TOKENS[i]` is followed by `TOKENS_AFTER_MAX_ACTIVATING_TOKEN[j]` in the text.Looking at the TOP_ACTIVATING_TEXTS:", bay leaves, chicken broth and water; bring to a boil then down to a simmer for 40 minutes. great challenge, go to, beautiful life, walk boldly

    np_acts-logits-general · gemini-2.5-flash-lite

    discussing options

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-pt/resid_post/layer_31_width_262k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     
    0.54
    2
    0.54
     och
    0.51
    Tra
    0.50
    )
    0.50
    Cat
    0.49
    Prediction
    0.49
    TR
    0.49
    -
    0.48
    Library
    0.46
    POSITIVE LOGITS
    yatiti
    0.56
    𝗴
    0.55
     अत्याचार
    0.55
    salaryfrom
    0.54
     সরাস
    0.54
    <unused37>
    0.54
    salaryto
    0.54
    ardı
    0.54
    swadian
    0.54
     больных
    0.53
    Activations Density 0.000%

    No Known Activations