© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 3-GEMMASCOPE-2-TRANSCODER-262K
    4. 91169
    Prev
    Next
    INDEX
    Explanations

    TOP_ACTIVATING_TEXTS:The activating texts mention:- Estonia (esti keel)- "unheimliche Geräusche von sich geben" (German: to make uncanny noises)- Germany (Rembrandt)- Middle East, purple murex dyes, royalty (Britain's Rulers, Celts)- Thule, sea, amber, Celts, Vikings- Celtic & Germanic Peoples, Romans, Goths, Vandals, Franks- Lions Mane Capsules, Nootropic, Brain Health- Cognitive enhancements, Prescription Medications, SupplementsLet's try to connect the lists.`MAX_ACTIVATING_TOKENS`: salt, ts, oot`TOKENS_AFTER_MAX_ACTIVATING_TOKEN`: ö, ame, amen, ?, who, **, were, ropic, ropThe most striking pattern is in `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` where "opic" appears after "oot". This strongly suggests "nootropic".Let's check if "nootropic" or related terms appear in `TOP_POSITIVE_LOGITS` and `TOP_ACTIVATING_TEXTS`.`TOP_POSITIVE_LOGITS`: ين, ك, 此, 1, ing, ре, ра, ة, الم, এThis list contains tokens in various languages (Arabic, Chinese, Russian, Bengali, Gujarati, Latin script). It doesn't immediately scream "nootropic". However, "ing" is a common English suffix, and "1" could be a placeholder or part of a number.`TOP_ACTIVATING_TEXTS`:- "Nootropic" is explicitly mentioned: - "Nootropic: (Targets users looking for cognitive benefits)" - "Supplements & Nootropics (External Boosts)"- Related concepts: "cognitive benefits", "cognitive enhancements", "Brain Health", "Brain Power".- "Lions Mane Capsules" - a common nootropic supplement.The pattern "oot" followed by "ropic" to form "nootropic" seems very strong given the activating texts."salt" followed by "ö" is less clear. Could it be part of "salto" or a corrupted word?Let's re-evaluate `MAX_ACTIVATING_TOKENS` and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` more broadly:- salt -> ö- ts -> ame, amen, ?, who, **, were (This could be related to plurals ending in 'ts', or roots like 'this', 'its', 'contents')- ts -> ame, amen, ?, who, **, were- ts -> ame, amen, ?, who, **, were- ts -> ame, amen, ?, who, **, were- ts -> ame, amen, ?, who, **, were- ts -> ame, amen, ?, who, **, were- oot -> ropic- oot -> ropic- oot -> ropicThe strongest signal is definitely `oot` + `ropic` -> `nootropic`.Since "nootropic" is a key term in the `TOP_ACTIVATING_TEXTS` and relates to cognitive enhancement, this is likely the primary focus of the neuron.How to phrase this concisely (3-20 words) without "tokens" or "patterns"?Initial thoughts:- nootropic supplements- cognitive enhancers- brain boosting- drugs for cognition- nootropic compounds- nootropic substancesLet's check the `TOP_POSITIVE_LOGITS` again. They seem very diverse, suggesting the neuron might be language-agnostic or activating on features that cross languages. "ing" is a common English suffix. "1" could be a number. "ре", "ра" could be parts of Slavic words. "ين", "ك", "ة", "الم" are Arabic. "此", "এ" are Chinese/Bengali.The neuron is definitely strongly associated with the concept of "nootropic".The phrase should capture what the neuron *detects* or *predicts*.Given the input:`MAX_ACTIVATING_TOKENS`: salt, ts, ts, ts, ts, ts, ts, oot, oot, oot`TOKENS_AFTER_MAX_ACTIVATING_TOKEN`: ö, ame, amen, ?, who, **, were, ropic, rop, ropThe most significant derivation is `oot` -> `nootropic`.The phrase should capture the essence of "nootropic".Let's consider the `TOP_POSITIVE_LOGITS`:These are words/tokens *associated* with the neuron (high positive logits).This list includes: ين, ك, 此, 1, ing, ре, ра, ة, الم, এThis list is very eclectic. "ing" is a general suffix. "1" could be a count. The others are from different scripts. It's hard to extract a meaning *directly* from this list that fits with "nootropic" *unless* some of these are parts of words that, when combined with other features the neuron detects, form concepts related to nootropics. However, the `MAX_ACTIVATING_TOKENS` and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` seem a more direct source of the *mechanism* or *trigger*.The primary, undeniable pattern is `oot` -> `ropic` forming `nootropic`.The neuron seems to be related to nootropics.Possible explanations:1. nootropics2. cognitive enhancement substances3. brain boosting supplements4. nootropic drug effects5. nootropic compounds"nootropics" is short and specific. It directly comes from `oot` + `ropic`.Let's check if "nootropic" itself appears in `TOP_POSITIVE_LOGITS`. No.So the neuron *finds patterns* that *lead to* nootropics.The phrase needs to be about what it *detects* or *predicts*.The most evident pattern derived is the formation of "nootropic".The `ACTIVATING_TEXTS` strongly reinforce this.What about `salt` -> `ö`? Or `ts` -> `ame`/`amen`/etc.?If the neuron detects "nootropic", it might also be sensitive to other related terms or word fragments that appear in similar contexts.The "ts" followed by "ame/amen/?" might relate to possession (its, ts) or plurals where 's' is added, but it's less clear.Focusing on the strongest signal: `oot` + `ropic` = `nootropic`.The explanation should be about these substances or their effects.Let's consider the prompt's examples:- "words related to", "concepts related to", "variations of the word" - avoid these.- "just say the pattern itself"- "could be about tokens following or preceding certain tokens"- "could be about words starting with a sequence"The most precise pattern derived from the tokens is the construction of the word "nootropic".So, the explanation should reflect this.Let's look at `TOP_POSITIVE_LOGITS` again. Does any part of it suggest nootropics in a broader sense?"ing" is a common suffix, present in "boosting", "enhancing", "improving". This is weak."1" could represent a count or a ranking.Given the strong presence of "nootropic" and related terms in the `

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_3_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    '
    1.27
    ’
    0.98
    ли
    0.85
     불구하고
    0.81
    <
    0.78
    <0x80>
    0.78
    (“
    0.78
    ");
    0.77
    ({\
    0.77
    				
    0.77
    POSITIVE LOGITS
    ين
    1.25
    ك
    1.08
    此
    0.98
    1
    0.91
    ing
    0.90
    ре
    0.87
    ра
    0.85
    ة
    0.84
    الم
    0.82
    এ
    0.81
    Activations Density 0.000%

    No Known Activations