© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-2-27B
    3. 22-GEMMASCOPE-RES-131K
    4. 33443
    Prev
    Next
    INDEX
    Explanations

    conspiracy accusations or refutations

    np_acts-logits-general · gemini-2.5-flash-lite

    The neuron responds to words that convey dramatic or sensationalized harm or excess—terms used to blame or describe extreme negative effects (e.g. “excessive,” “noise,” “causing,” “accused,” “earthquakes”).

    oai_token-act-pair · o4-miniTriggered by @jyhe0408
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-27b-pt-res/layer_22/width_131k
    Prompts (Dashboard)
    24,576 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     Touristen
    -0.96
    clud
    -0.94
     remedio
    -0.90
     jesús
    -0.90
     garcía
    -0.90
    ]+
    -0.86
    ebly
    -0.85
     ""){
    -0.85
    公斤
    -0.85
     propone
    -0.83
    POSITIVE LOGITS
     imminent
    1.03
    Comparisons
    0.92
    べて
    0.92
    bolso
    0.92
     doom
    0.90
     deswegen
    0.87
     entweder
    0.86
    oine
    0.86
    canes
    0.85
     '@/
    0.85
    Activations Density 0.040%

    No Known Activations