© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-12B-IT
    3. 12-GEMMASCOPE-2-RES-16K
    4. 12640
    Prev
    Next
    INDEX
    Explanations

    I need to find a pattern that explains what this neuron detects or predicts, using the provided lists.1. **MAX_ACTIVATING_TOKENS**: All tokens are "to". This is a very strong indicator.2. **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: The tokens immediately following "to" are `choose`, `the`, `know`, `be`, `the`, `the`, `work`, `a`, `the`, `work`. Looking at these, we see common constructions like "to choose", "to know", "to be", "to work". These are infinitive verb forms. `the` and `a` are articles that can appear after "to" in specific contexts, but `choose`, `know`, `be`, `work` are verbs.3. **TOP_POSITIVE_LOGITS**: `م`, `ين`, `та`, `зна`, `ال`, `Estados`, `Só`, `Lista`, `جاد`, `м`. These are tokens from various languages.infinitive verbs

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-12b-it/resid_post/layer_12_width_16k_l0_medium
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    り
    1.89
    서
    1.77
     a
    1.68
     এছাড়া
    1.62
    ism
    1.48
     sped
    1.48
     suspiciously
    1.48
     Canc
    1.47
    ").
    1.46
    up
    1.44
    POSITIVE LOGITS
    م
    1.93
    ين
    1.80
    та
    1.78
    зна
    1.72
    ال
    1.69
    Estados
    1.65
    Só
    1.63
    Lista
    1.62
    جاد
    1.58
    м
    1.55
    Activations Density 0.034%

    No Known Activations