© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B
    3. 31-GEMMASCOPE-2-RES-262K
    4. 671
    Prev
    Next
    INDEX
    Explanations

    military matters (forces, war)2. historical entities (Khan)3. institutions (school)These often lead to subsequent events or changes in state, indicated by `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` and `TOP_POSITIVE_LOGITS`.Looking at `MAX_ACTIVATING_TOKENS`, `forces`, `Khan`, `School`, `War` are prominent.Looking at `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`, we see `prepared`, `retreated`, `again`, `under`.Looking at `TOP_POSITIVE_LOGITS`, we see `soon`, `thereafter`, `became`, `eventually`, `subsequently`, `began`, `later`.The pattern seems to be about entities (forces, Khan, school) and events (War) that have consequences or happen before/after other events.Let's try to synthesize a phrase."forces" + "Khan" + "school" + "War" suggests specific types of entities or contexts.The subsequent words and logits suggest progression or outcomes."Khan forces prepared""forces retreated""school under""Khan retreated"The logits like "became", "later", "subsequently" strongly imply a temporal or consequential relationship.The texts talk about historical events, military campaigns, and institutional changes.Let's consider the most distinctive elements in MAX_ACTIVATING_TOKENS: `forces`, `Khan`, `school`, `War`. These are specific nouns.The tokens after often describe actions or states: `prepared`, `retreated`, `under`, `was`.The positive logits are temporal: `later`, `subsequently`.So, it's about entities involved in events, and what happens later."forces Khan school war later" - Too literal."entities involved in later events" - Too abstract."military forces and leaders" - Misses school."historical events and participants" - Getting closer."forces, leaders, institutions, then consequences" - Too long.Let's look at the core entities: `forces`, `Khan` (a ruler/leader), `school` (an institution). `War` is an event.The time words (`later`, `subsequently`, `became`) are key.Could it be about past military or historical events and their outcomes?"military and historical events aftermath" - 5 words."historical events war outcomes" - 4 words."forces, war, then subsequent actions" - 6 words.Let's rethink the MAX_ACTIVATING_TOKENS: `before`, `forces`, `Khan`, `but`, `he`, `aj`, `school`, `War`, `,`, `school`.The key words are `forces`, `Khan`, `school`, `War`, `before` (temporal).The TOKENS_AFTER_MAX_ACTIVATING_TOKEN are `prepared`, `retreated`, `was`, `again`, `under`. These are verbs/prepositions indicating action, state change, or consequence.The TOP_POSITIVE_LOGITS are explicitly temporal/sequential: `soon`, `thereafter`, `became`, `eventually`, `subsequently`, `began`, `later`, `thence`.This neuron is strongly tracking historical events involving military forces and leaders, and what happened *after* them. historical events and their aftermath

    np_acts-logits-general · gemini-2.5-flash-lite

    historical events and figures

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-pt/resid_post/layer_31_width_262k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     naprawdę
    0.44
    我们的
    0.43
    我們的
    0.43
     potencialmente
    0.42
     mūsų
    0.41
     nossa
    0.40
     هستند
    0.40
     funktioniert
    0.40
     yüzde
    0.40
     naszego
    0.40
    POSITIVE LOGITS
     soon
    0.45
     thereafter
    0.45
     became
    0.43
     eventually
    0.41
     subsequently
    0.40
    became
    0.40
     began
    0.40
     erneut
    0.39
     later
    0.38
     thence
    0.38
    Activations Density 0.191%

    No Known Activations