© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 31-GEMMASCOPE-2-RES-262K
    4. 82353
    Prev
    Next
    INDEX
    Explanations

    threats and violence

    np_acts-logits-general · gemini-2.5-flash-lite

    sentences where the model talks about itself, its capabilities, training, or safety/limitation explanations (meta/self-referential assistant statements).

    oai_token-act-pair · gpt-5-miniTriggered by @yooniel31
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/resid_post/layer_31_width_262k_l0_medium
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    дка
    0.48
    桢
    0.46
    ጡ
    0.45
    ଐ
    0.43
    ဃ
    0.43
    脸
    0.43
    મા
    0.42
    状态
    0.42
    uten
    0.41
    ста
    0.41
    POSITIVE LOGITS
     authorizes
    0.44
    canic
    0.43
    ോ
    0.43
    👏👏
    0.42
     alleviate
    0.41
     gunfire
    0.41
    GÁ
    0.41
    ristmas
    0.40
     authorizing
    0.39
     മുഴ
    0.39
    Activations Density 0.004%

    No Known Activations