© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Qwen3-1.7B
    3. 26-LLAMASCOPE-2-LORSA-16K-K64
    4. 1247
    Prev
    Next
    INDEX
    Explanations

    <THINKING>Method 1 works: MAX_ACTIVATING_TOKENS repeatedly show self‑attention patterns where the source and target token are “anger” (e.g., [(., joy,) anger] => ( joy, anger)). This indicates the neuron activates on the word “anger”. Other methods give no clear consistent pattern. Therefore, the neuron detects “anger”. Method 1 used.</THINKING> anger

    unknown · unknown
    New Auto-Interp
    Top Features by Cosine Similarity
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     anger
    -30.63
     angi
    -22.75
     angl
    -22.63
     Ang
    -22.38
    Eng
    -22.38
    _ang
    -22.25
     angry
    -22.00
    eng
    -22.00
    Ang
    -21.13
     angular
    -21.13
    POSITIVE LOGITS
     economic
    20.88
     Economic
    20.63
    经济
    20.50
    -economic
    19.38
    /script
    18.63
    經濟
    18.50
     Palin
    18.38
     pal
    18.38
    (script
    18.13
     SCRIPT
    17.00
    Activations Density 0.103%

    No Known Activations