© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Olmo-3-1125-32B
    3. 32-RES-BATCHTOPK-131K
    4. 91078
    Prev
    Next
    INDEX
    Explanations

    Analyzing the provided lists:* **MAX_ACTIVATING_TOKENS**: rat, sat, Nazi, Garcia, Trump, Cobra, appellant, ery, ons, :* **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: emotionally, an, officer, 's, as, ,, ,, NVIDIA, of, opencv* **TOP_POSITIVE_LOGITS**: greed, arrog, disgustinrg, bully, brag, notorious, coward, selfish, abusive, greedy* **TOP_ACTIVATING_TEXTS**: Contains examples like "satan's storage unit", "female Nazi officer", "Donald Trump", "Cobra, have become an international peacekeeping force", "appellant, for personal injury".The "TOP_POSITIVE_LOGITS" heavily suggest negative traits or concepts.The "MAX_ACTIVATING_TOKENS" include names of entities (Nazi, Trump, Cobra, Garcia) and specific roles/statuses (appellant, officer).The "TOP_ACTIVATING_TEXTS" provide contexts where these entities or roles are discussed, often in negative or contentious situations (e.g., "emotionally blackmails", "satan's", "Nazi officer", "sadistic homosexual mercenary", "Trump as a candidate", "enemies, Cobra", "personal injury sustained by appellee", "defective panel raiser").The pattern seems to be related to specific entities, individuals, or groups often associated with negative contexts, conflict, or specific domains. The logit list strongly points towards negative descriptors.Therefore, the neuron seems to be activated by mentions of distinct entities or roles that are frequently contexted negatively or with significant implications.Explanation: **notable figures and entities in conflict**Let me refine this to be more direct and concise as per the rules.Looking at the TOP_POSITIVE_LOGITS, they are all negative descriptors.Looking at MAX_ACTIVATING_TOKENS, we see named entities or specific roles (Nazi, Garcia, Trump, Cobra, appellant).The TOP_ACTIVATING_TEXTS show these entities in contexts that are often negative or consequential (e.g., "emotionally blackmails", "satan's", "Nazi officer", "mercenary", "Trump as a candidate", "enemies

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    bcywinski/Olmo-3-32B-Base-SAE/saes_allenai_Olmo-3-1125-32B_batch_top_k/resid_post_layer_32
    Prompts (Dashboard)
    16,384 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     heroic
    -0.11
     bona
    -0.10
     mild
    -0.10
     bless
    -0.10
     brave
    -0.10
    809
    -0.10
     admirable
    -0.10
     allot
    -0.10
     mat
    -0.10
     gently
    -0.10
    POSITIVE LOGITS
     greed
    0.17
     arrog
    0.16
     disgusting
    0.16
     bully
    0.16
     brag
    0.16
     notorious
    0.15
     coward
    0.15
     selfish
    0.15
     abusive
    0.14
     greedy
    0.14
    Activations Density 0.309%

    No Known Activations