© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 9-GEMMASCOPE-2-TRANSCODER-262K
    4. 194464
    Prev
    Next
    INDEX
    Explanations

    I need to provide a concise explanation (3-20 words) about what the neuron detects or predicts, based on the provided lists.1. **MAX_ACTIVATING_TOKENS**: The token "stats" appears repeatedly and exclusively. This is a strong indicator.2. **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: Frequent tokens are "import" and "as". This suggests a pattern like "stats import" or "stats as".3. **TOP_POSITIVE_LOGITS**: This list contains non-English characters and abbreviations. It's harder to find a direct, consistent pattern here that strongly complements the others, but it might suggest regional or specific technical vocabulary related to stats.4. **TOP_ACTIVATING_TEXTS**: These texts are overwhelmingly code snippets, specifically Python code, mentioning `pandas`, `scipy.stats`, `chi2_contingency`, `numpy`, `matplotlib`, `seaborn`, `pointbiserialr`, `spearmanr`, `gamma`, `norm`, `ks_2samp`, `pearsonr`, `skew`, `kurtosis`. The common theme is statistical analysis and libraries, often imported.Combining these:- "stats" is the key token.- It's often followed by "import" or "as".Detecting "stats import" sequences

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_9_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     trusses
    0.54
     Jimi
    0.53
    '}
    0.52
    เรื่อง
    0.51
    ढ
    0.51
     I
    0.50
    '`
    0.50
     Downey
    0.50
     Opie
    0.50
     subconsciously
    0.50
    POSITIVE LOGITS
    är
    0.59
    é
    0.55
    вър
    0.54
    á
    0.52
    AT
    0.52
    AL
    0.51
    važ
    0.50
     때
    0.49
     orientación
    0.49
     و
    0.49
    Activations Density 0.000%

    No Known Activations