© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 32-GEMMASCOPE-2-TRANSCODER-262K
    4. 235924
    Prev
    Next
    INDEX
    Explanations

    `**TOP_ACTIVATING_TEXTS**: * "attempts to overturn the 2020 election" * "predicted reckoning" * "leading to widespread protests and" * "stopping the corruption" * "alter an outcome" * "opposed to Trump" * "questioning the election."**Pattern Identification:*** **MAX_ACTIVATING_TOKENS** and **TOKENS_AFTER_MAX_ACTIVATING_TOKEN** show a strong pattern around actions following "to": "to overturn", "to alter", "to stopping".* **TOP_ACTIVATING_TEXTS** heavily features themes of: * "overturn the 2020 election" * "predicted reckoning" * "widespread protests" * "stopping corruption" * "alter an outcome" * "opposed to Trump" * "questioning the election"* The combination points towards actions related to challenging, overturning, or altering election outcomes and significant societal/political events, often involving "Trump" and "corruption" or calls for "stopping" something.* The **TOP_POSITIVE_LOGITS** list seems largely irrelevant due to its random nature and non-English characters, except for perhaps "YEARS" and "WHY" which are too generic.**Synthesizing the Pattern:**The neuron seems to activate when texts discuss challenging existing states, outcomes, or authority, particularly concerning elections and corruption, or predicting significant future events like a "reckoning." The action verbs associated with these challenges are key.**Candidate Explanations:**1. overturning, altering, stopping actions2. election disruption and reckoning3. challenging outcomes and corruption4. predicting reckoning and overturning5. actions to overturn election or stop corruptionLet's refine based on conciseness and specificity. The core theme is *action verbs* indicating a challenge or disruption, especially in political/election contexts."overturning, altering, stopping actions" is very direct."election disruption and reckoning" captures the context well."challenging outcomes and corruption" is also good.Considering the specific examples like "overturn the 2020 election", "predicted reckoning", and "stopping the corruption", a phrase that encompasses these actions and predictions is ideal."overturning election outcomes and reckoning""challenging elections, corruption, and predicting reckoning""actions to overturn, alter, stopping corruption"The prompt asks for what the neuron *detects or predicts*.The texts strongly suggest election-related challenges and future predictions."overturn election outcomes and predictions" - This is 5 words."election challenges, corruption, and reckoning" - 5 words."actions to overturn election and stop corruption" - 7 words.Let's look at the `MAX_ACTIVATING_TOKENS` for emphasis again.'to' -> 'overturn''to' -> 'overturn''predicted' -> 'reckoning''to' -> 'widespread' (used in context of protests)'stopping' -> 'legal''and' -> 'alter''Trump''might'The pattern is about performing actions that change or challenge existing states. "Overturn", "alter", "stopping", "reckoning" are key indicators."overturning, altering, stopping and reckoning" - 5 words. This captures the verbs and the prediction."challenging election outcomes and predicting reckoning" - 6 words. This is more specific to the themes observed."actions to overturn, alter, stop corruption" - 7 words. This is also quite specific.Let's go with a phrase that highlights the nature of the action.Final check: 3-20 words. Specific. No "This neuron detects". No "tokens". No capitalization unless proper noun."overturning outcomes and predicting reckoning" (5 words) - This seems to cover the core pattern from the texts and tokens well. "Outcomes" covers elections and other results. "Predicting reckoning" covers the future event prediction.Another option: "actions to overturn and stop corruption" (6 words). This is also very strong.Comparing:"overturning outcomes and predicting reckoning""actions to overturn and stop corruption"The list also has "alter" and "widespread". The first option seems broader."predicted reckoning", "overturn the 2020 election", "stopping the corruption", "alter an outcome", "widespread protests".Let's try to combine the *action* aspect with the *target*."challenging election outcomes and corruption" (5 words)"overturning elections and predicting reckoning" (5 words)The word "predicted" is in MAX_ACTIVATING_TOKENS. The TOKENS_AFTER include "reckoning", "widespread", "alter", "Trump". The Texts include "overturn", "reckoning", "widespread protests", "stopping corruption", "alter an outcome", "Trump".The pattern is about actions that disrupt or change a state (overturn, alter, stop) and predicted future events (reckoning)."overturning, altering, stopping, and reckoning" - 5 words. Excellent.overturning, altering, stopping, and reckoning

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_32_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     ubiquitous
    0.41
     overriding
    0.39
    attering
    0.39
     terbuka
    0.38
     overlying
    0.38
     punctatis
    0.36
    χη
    0.36
     protest
    0.35
    boolProp
    0.35
    риб
    0.35
    POSITIVE LOGITS
     कसा
    0.41
     космо
    0.39
     YEARS
    0.39
    ソコン
    0.38
     Bavaria
    0.38
     WHY
    0.38
     कृति
    0.37
    ueless
    0.37
    Alessandro
    0.37
    ెస్
    0.36
    Activations Density 0.000%

    No Known Activations