© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 6-GEMMASCOPE-2-TRANSCODER-262K
    4. 133101
    Prev
    Next
    INDEX
    Explanations

    The neuron seems to activate on or around specific punctuation, often followed by proper nouns or thematic keywords that typically introduce a synopsis or a named entity.1. **MAX_ACTIVATING_TOKENS**: Contains punctuation like `"`, `“`, `)`, `*`, `:`, `**`, and empty strings. This suggests the neuron triggers on structural markers or delimiters.2. **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: Shows tokens like `A`, `Arthur`, `Dust`, `The`, `Kai`, `New`, `Billion`. These are often the start of sentences/phrases or names.3. **TOP_POSITIVE_LOGITS**: Contains Arabic characters (`نا`, `ول`, `ية`, `لی`, `انا`) and numbers/English tokens (`2`, `6`, `jons`, `र्ड`). This suggests a mix of language or potentially very specific code/numeric patterns, but the others are more indicative.4."synopsis" delimiter or special character

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_6_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    a
    0.73
    ه
    0.67
    ה
    0.62
     plupart
    0.55
     progett
    0.52
     mauvaise
    0.52
    н
    0.50
    И
    0.47
    К
    0.47
     scoff
    0.46
    POSITIVE LOGITS
    نا
    0.50
    <
    0.48
    ول
    0.48
    2
    0.42
    6
    0.42
    jons
    0.41
    ية
    0.41
    لی
    0.40
    र्ड
    0.40
    انا
    0.39
    Activations Density 0.001%

    No Known Activations