© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 31-GEMMASCOPE-2-TRANSCODER-262K
    4. 163149
    Prev
    Next
    INDEX
    Explanations

    This neuron is identifying occurrences of negative sentiment or uncertainty, often related to actions, states, or desires, particularly when described in relation to female subjects. It seems to capture phrases where something is being stated as not happening, not being, or not being wanted, or where uncertainty is expressed.TOP_POSITIVE_LOGITS: OHAMA, Watts, ோம், 쾌, respir, xspace, শিশ,OrFail, AMOS, なくなる * These are diverse, some Chinese, Japanese, Bengali, and English names/words. * They don't immediately reveal a clear linguistic pattern related to the other lists. * 'なくなる' (nakunaru) in Japanese means "to disappear" or "to be lost". This aligns with a negative/loss sentiment.4. **TOP_ACTIVATING_TEXTS**: * 'seems eager to end the conversation, he might **not be interested**.' (negation, interest) * 'someone might **not be direct**.' (negation, directness) * 'If she's pulling away, crossing her **arms**, or avoiding eye contact, *pump the brakes*.' (physical cues, implying negative feelings/lack of interest) * 'alternating between showing interest ("hot") and withdrawing it ("cold"). The PUA might be intensely complimentary one moment, then dismissive or aloof' (withdrawing interest, dismissive, aloof - all negative/lack of interest) * '"I don't want to see you anymore, POLICE"' (strong negation, rejection) * 'They will *absolutely* **not work** on everyone.' (negation,

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_31_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     mutual
    0.34
     czerw
    0.33
     buttons
    0.32
     červ
    0.31
     chaperone
    0.30
     patronage
    0.30
     Gesture
    0.30
     Mutual
    0.30
     marker
    0.29
     clicking
    0.29
    POSITIVE LOGITS
    OHAMA
    0.30
    Watts
    0.30
    ோம்
    0.29
    쾌
    0.29
     respir
    0.29
    ianSpace
    0.28
     শিশ
    0.28
    OrFail
    0.27
    AMOS
    0.27
    なくなる
    0.27
    Activations Density 0.015%

    No Known Activations