© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-2-27B
    3. 10-GEMMASCOPE-RES-131K
    4. 68176
    Prev
    Next
    INDEX
    Explanations

    thought processThe user wants me to describe the neuron's behavior in 3-20 words, based on the provided lists, focusing on detected patterns.1. **Analyze `MAX_ACTIVATING_TOKENS`**: - `knew` - `thought` - `decided` - `wondering` - `think` - `tell` - `imagine` These tokens clearly relate to mental states, cognition, judgment, or understanding.2. **Analyze `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`**: - `there` - `to` - `they` - `“` - `which` - `where` - `it` - `of` - `me` - `as` - `the` The presence of quotation marks (`“`) is significant. This suggests that the cognitive tokens from `MAX_ACTIVATING_TOKENS` are often followed by spoken or written declarations, opinions, or questions.3. **Analyze `TOP_ACTIVATING_TEXTS`**: - "thought there is a good chance..." - "thought to myself “that looks like a good idea”." (Max activating token 'thought' followed by quote) - "knew they were destined..." - "thought “these aren’t going to do anything”..." (Max activating token 'thought' followed by quote) - "decided which was right..." - "wondering where their electrical rates are going." - "knew it was right up my alley..." - "think of me." - "tell me where I made the mistake ?" - "knew there was no way I was walking without them." - "knew the minds of their men" - "imagine as you..." The texts confirm the interpretation of `MAX_ACTIVATING_TOKENS` as cognitive states. They also show that after these states, there's often a specific outcome, opinion, or a direct quote/question.4. **Analyze `TOP_POSITIVE_LOGITS`**: - `그것` (Korean for "it") - `內` (Chinese/Japanese for "inside") - `étan` (French, related to "being") - `écrire` (French for "to write") - `(゚` (Emoticon) - `腳` (Chinese/Japanese for "foot") - `帶` (Chinese/Japanese for "bring") - `Regarding` (English) - `貓` (Chinese/Japanese for "cat") - `expédition` (French for "expedition") This list is quite diverse and includes many non-English words. While specific words don't form a direct semantic link with "thought" or "knew", the *diversity* and inclusion of words like "écrire" (to write), or the implied reporting/statement nature ("Regarding", "étan"), might hint at the neuron firing when something is being stated, decided upon, or written down. The foreign words suggest the neuron is not language-specific in its core function, but rather captures a *type* of cognitive action that can occur in any language.5. **Synthesize a pattern**: The core function seems to be about internal cognitive processes that lead to a decision, judgment, or statement. The phrase pattern is often "cognitive verb + outcome/statement". Given the `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` include `“`, and `TOP_ACTIVATING_TEXTS` show examples like "thought '...',", the neuron is strongly associated with expressing these cognitive moments. Let's try to combine these: - "Cognitive processes leading to a statement" (Too long, slightly descriptive) - "Thinking followed by a declaration" (Good, fits the pattern) - "Making a judgment and stating it" (Good) - "Decisions and opinions expressed" (Good) - "Thoughts leading to conclusions" (Okay, but less about the *expression*) - "Stating a thought or decision" (Concise and fits) Considering `MAX_ACTIVATING_TOKENS` (knew, thought, decided, wonder, think, tell, imagine) and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` (especially `“`), a phrase like "stating a thought or decision" or "expressing judgments and beliefs" works well. "Stating a thought or decision" is more direct. Let's check constraints: 3-20 words. "stating a thought or decision" is 5 words. - No "detects/predicts". Check. - No "tokens" or "patterns". Check. - Proper noun capitalization. None needed. Check. - Specificity. "stating a thought or decision" is specific enough. Check. - Majority match. The evidence strongly supports this. Check. Another angle: The presence of foreign words in `TOP_POSITIVE_LOGITS` and multilingual `TOP_ACTIVATING_TEXTS` suggests the neuron is cross-lingual. The pattern isn't just about English words. The core is the *act* of cognitive formulation and declaration. Could it be simpler? "Expressing a mental state" - A bit abstract. "Formulating an opinion" - Good, but "decided" and "knew" imply more than just opinion. "Internal decision making and outward expression" - Too long. Let's re-evaluate the combination of MAX_ACTIVATING_TOKENS and TOKENS_AFTER_MAX_ACTIVATING_TOKEN. - `knew` + `there` - `thought` + `to`

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-27b-pt-res/layer_10/width_131k
    Prompts (Dashboard)
    24,576 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     visit
    -1.22
     include
    -1.21
     on
    -1.21
     you
    -1.14
    ↵↵↵↵
    -1.11
     behind
    -1.10
    3
    -1.09
     membuat
    -1.08
    ↵↵↵
    -1.07
     into
    -1.06
    POSITIVE LOGITS
     그것
    1.43
     內
    1.34
     étan
    1.30
    écrire
    1.29
    (゚
    1.25
     腳
    1.22
     帶
    1.22
    Regarding
    1.20
     貓
    1.20
    expédition
    1.19
    Activations Density 0.112%

    No Known Activations