© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Qwen3-32B
    3. 32-RESID-BATCHTOPK-65K
    4. 2922
    Prev
    Next
    INDEX
    Explanations

    thought processThe user wants me to explain the behavior of a neuron based on the provided lists.The explanation needs to be a concise phrase (3-20 words) describing a pattern, not a full sentence, and avoid specific introductory phrases.1. **Analyze `MAX_ACTIVATING_TOKENS`**: `2`, `smart`, `conversations`, `thinker`, `bright`, `challenge`, `opponents`, `!` These tokens strongly suggest concepts related to intelligence, thinking, discussion, and debate.2. **Analyze `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`**: `0`, `out`, `and`, `.`, `,`, `.`, `,`, `received`, `by`, `The` These are mostly neutral or connecting tokens, not offering a strong pattern related back to the `MAX_ACTIVATING_TOKENS`.3. **Analyze `TOP_POSITIVE_LOGITS`**: `辩论`, `æįįåį«`, `论è¯ģ`, `pont`, `驳`, `defending`, `羣çIJĨ`, `辩`, `æ»Ķ`, `äºī论` The majority of these are Chinese/Japanese characters which, when translated or interpreted, often relate to discussion, logic, opinion, or argument. Of particular note are `论è¯ģ` (argumentation/discourse), `defending` (which is also in English logits).thought, debate, and arguments

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    adamkarvonen/qwen3-32b-saes/saes_Qwen_Qwen3-32B_batch_top_k/resid_post_layer_32
    Prompts (Dashboard)
    16,384 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     Melania
    -0.09
    æļĹ示
    -0.09
    å±Ģéķ¿
    -0.09
    æĢľ
    -0.09
    _imag
    -0.08
    ween
    -0.08
    çĹħ
    -0.08
    æĺ¥
    -0.08
    éģ¿
    -0.08
    å¾·æĭī
    -0.08
    POSITIVE LOGITS
    辩论
    0.15
    æįįåį«
    0.11
    论è¯ģ
    0.11
     pont
    0.10
    驳
    0.10
     defending
    0.10
    羣çIJĨ
    0.10
    辩
    0.10
    æ»Ķ
    0.10
    äºī论
    0.09
    Activations Density 0.292%

    No Known Activations