© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 40-GEMMASCOPE-2-RES-262K
    4. 193856
    Prev
    Next
    INDEX
    Explanations

    course networks weight shifts scientific USB---* **course** `that` (Develop a curriculum for a high school **course** `that` teaches students...)* **olds** `are` (1-year-**olds** `are` tough on toys!)* **personal** `networks` (business, marketing, and potentially even **personal** `networks`.)* **weight** `shifts` (the **weight** `shifts` considerably throughout different stages of life.)* **potential** `futures` (reaction *to* **potential** `futures`, not a prediction of them.)* **weight** `calorie` (Alcohol is calorie-dense and can hinder **weight** loss.) - *This one has 'calorie' after 'weight', which fits the pattern better than 'shifts'. The neuron might associate 'weight' with nutrition/loss.** **scientific** `methods` (**scientific** `methods`, project management workflows)* **USB** `input` (A **USB** Mouse and Keyboard: Some TVs support **USB** `input`, but compatibility varies.)* **Generating** `a` (comment or description for 'Generating a random integer', fitting the idea of generative processes).The pattern seems to be nouns or adjectives that are often followed by specific types of words (often nouns or verbs indicating a state, action, or category).Let's try to combine "course", "networks", "weight", "scientific", "USB", and consider their typical follow-ups:- course -> that, for, in- olds -> are, and- personal -> networks, relationships, finance- weight -> shifts, loss, gain, management, calorie- potential -> futures, risks, benefits, clients- scientific -> methods, research, community, papers- USB -> input, drive, port, cableThe `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` list is key here.- course -> **that**- olds -> **are**- personal -> **networks**- weight -> **shifts** (or **calorie** from the text)- potential -> **futures**- scientific -> **methods**- USB -> **input**The pattern is often a specific noun/concept followed by another noun/verb associated with it (like 'course that', 'personal networks', 'weight shifts', 'scientific methods', 'USB input').The common theme links concepts to their subsequent categories or descriptions.Given the `TOP_POSITIVE_LOGITS` list which contains tokens like "텟", "없이", ")/", "splitext", "generar", "Bola", "vasena", these seem unrelated to the English words found in `MAX_ACTIVATING_TOKENS` and `TOP_ACTIVATING_TEXTS`. This suggests the neuron might be bilingual or have learned patterns across languages, but the user prompt asks to focus on patterns in the *lists provided*. Since the English words are clearer and have direct connections in the `TOP_ACTIVATING_TEXTS`, I will focus on those.The pattern is "concept + related term"."course networks weight scientific USB" are all concepts that appear frequently. The words following them in `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` refine what these concepts are related to.Let's try to find a concise phrase that captures this:- "specific concepts followed by related nouns" -> too long- "concept + descriptor" -> 3 words, good.Looking at the examples:* "course that"* "personal networks"* "weight shifts"* "scientific methods"* "USB input"This seems to be the core behavior. The neuron activates when it sees these specific concepts followed by terms that describe their context, purpose, or consequence.The phrase "concept + descriptor" or "conceptual association" fits."concept + descriptor" is 3 words."conceptual association" is 2 words. The prompt asks for 3-20 words."finding context for concepts" -> 4 words.Let's re-examine the prompt: "concise explanation (3 to 20 words) that captures what the neuron detects or predicts by finding patterns in lists."The lists demonstrate the neuron detecting specific concepts and then predicting or associating them with subsequent related terms.Consider the most direct way to state this pattern."concept followed by related term" - 4 words. This is accurate."specific terms with their continuations" - 4 words.Let's check the provided `MAX_ACTIVATING_TOKENS` and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` again:- course -> that- olds -> are- personal -> networks- weight -> shifts / calorie- potential -> futures- While -> (no clear short pattern word after it)- scientific -> methods- USB -> input- you -> ' (often precedes punctuation)- . -> (end of sentence indicator)The strongest pattern is noun/adjective followed by another noun/verb that elaborates on it.I should avoid stating "tokens" or "patterns"."concepts and their contexts" - 3 words. This seems to capture the essence."associating concepts with follow-up words" - 4 words."specific concepts and subsequent words" - 4 words.Let's pick the simplest and most direct. "concepts and their contexts" feels good.Let's evaluate it against other possibilities:* "concept + descriptor" - too short, 3 words. "concept" and "descriptor" are meta-terms.* "associating terms and their follow-ups" - 5 words.* "identifying concepts and their continuations" - 5 words.The core behavior is predicting *what comes next* or *what is associated with* a given term.The output of `TOP_POSITIVE_LOGITS` appears to be Chinese/other non-English characters. This might mean the neuron is multilingual or the pattern is abstract. However, the clear English pattern in `MAX_ACTIVATING_

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/resid_post/layer_40_width_262k_l0_medium
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    यर
    0.37
    >∈</
    0.37
    傻
    0.36
    right
    0.35
    一会儿
    0.35
    game
    0.35
     гаран
    0.35
     inoxyd
    0.35
    Label
    0.35
    creative
    0.34
    POSITIVE LOGITS
    텟
    0.44
     없이
    0.42
    ")/
    0.40
     fent
    0.40
    }$.)
    0.40
    splitext
    0.40
    wała
    0.40
     generar
    0.39
    Bola
    0.39
    vasena
    0.38
    Activations Density 0.001%

    No Known Activations