© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 12-GEMMASCOPE-2-TRANSCODER-262K
    4. 202383
    Prev
    Next
    INDEX
    Explanations

    - "dem" + "elles" -> "demoiselles" (from Picasso text)- "demi" + "plane" -> "demiplane" (from Ravenloft text)- "ém" + "ocrat" -> "émocrate" (French for democrat, context is a political party)- "dem" + "ilitar" -> "demilitarized" (from Israeli settlements text)- "lord" + ";" -> seems like a sentence ender, context is "demon lord"The tokens "dem" and "demi" are prominent in `MAX_ACTIVATING_TOKENS`.The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` list shows completions like "elles", "plane", "ilitar", "ocrat".Combining "dem" with "elles" gives "demoiselles".Combining "demi" with "plane" gives "demiplane"."lord" appears followed by a separator."ém" + "ocrat" appears to be part of "démocrate"."dem" + "ilitar" appears to be part of "demilitarized".The common theme seems to be words or parts of words that relate to social, political, or fictional hierarchical structures, often with specific suffixes.Looking at `TOP_POSITIVE_LOGITS`: SignIn, Squares, DEFGHIJKLMNOP, Gard, Incre, signUp, ेशन, sing, समिट, Health. This list is diverse and doesn't immediately show a strong pattern with the activating tokens, except perhaps a general sense of naming or classification.However, `TOP_ACTIVATING_TEXTS` is crucial:- "Les Demoiselles d'Avignon" -> "Demoiselles"- "demiplane" -> "demiplane"- "SMER-SD (Direction – Social Démocratie)" -> "Démocratie"- "demilitarized" -> "demilitarized"- "demon lord" -> "lord"The neuron is strongly activating on sequences related to "dem", "demi", "démocra", "demoiselles", "demilitarized", and "lord".It seems to be detecting words related to specific social/political/fictional entities or concepts.A strong candidate is the prefix "dem" or "demi", and completions related to states or titles."Demoiselles", "demilitarized", "demiplane", "démocratie".Let's refine the pattern. It's picking up these specific word roots and often completing them.It seems to be capturing "dem" + specific endings."dem" + "oiselles""demi" + "plane""dem" + "ilitarized" (implied)"démocra" + "tie"The neuron seems to activate for words starting with "dem" or "demi" and often followed by specific semantic endings.Let's look for a concise phrase:- "dem" followed by specific suffixes- words with "dem" prefix- "dem" and "demi" related conceptsThe `TOP_POSITIVE_LOGITS` like "SignIn" and "signUp" suggest actions or entities, but not directly linked to "dem". "Gard" could be related to "guard" or names. "Sumit" is "summit". "Health" is generic. "DEFGHIJKLMNOP" is alphabets. "Incre" is prefix for increase.The core pattern is definitely in the `MAX_ACTIVATING_TOKENS` and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` in relation to `TOP_ACTIVATING_TEXTS`.It's about words starting with "dem" or "demi", often referring to entities, fictional worlds, political/social states, or descriptive terms.Consider the completions:"dem" -> "oiselles", "ilitar""demi" -> "plane""ém" -> "ocrat""osc" -> "ener""lord" -> ";" (end of sentence)The commonality is the prefix "dem" or "demi" leading to distinct concepts."Demoiselles" (art), "demilitarized" (military/political), "demiplane" (fiction/gaming), "démocratie" (political). "Lord" is a title.This neuron seems to be identifying complex noun phrases or concepts that often start with "dem-" or relate to specific roles/titles like "lord".A concise way to capture this:- "dem" and "lord" related concepts- words from "demoiselles" to "demilitarized"- entities/states named with "dem-"The phrase should capture the essence of these words."demoiselles", "demilitarized", "demiplane", "démocratie", "lord".These are specific concepts.How about focusing on the root and the type of words it leads to?- "dem" prefix words. (Too broad)- "dem" and "demi" word completions. (A bit literal)- Specific "dem" and "lord" concepts. (Getting closer)Let's look at the `TOP_POSITIVE_LOGITS` again. They are quite varied. The neuron might be contributing to identifying specific *types* of named entities or concepts.The most distinctive commonality is the "dem" prefix, often leading to specific contexts.- "Demoiselles" (art)- "Demilitarized" (politics/military)- "Demiplane" (fiction)- "Démocratie" (politics)- "Lord" (title)The neuron is very specific. It's not just any "dem" word.Could it be about titles, specific entities, or states?The *type* of word is important.Trying to find a pattern that covers "demoiselles", "demilitarized", "demiplane", "démocratie", "lord".These are all nouns or descriptors for specific things.Let's consider the first few MAX_ACTIVATING_TOKENS and their completions:"dem" + "elles""demi" + "plane""lord" + ";"This strongly suggests words that start with "dem" and have specific continuations forming established terms or concepts.Let's consider the prompt's examples:- "finding patterns in lists."- "what the neuron detects or predicts"- "concise explanation (3 to 20 words)"- "single word, or phrase, or pattern."- "explanation could be about tokens following or preceding certain tokens."- "explanation could be about words starting with a sequence."- "do not mention 'tokens' or 'patterns' in your explanation."Given "demoiselles", "demilitarized", "demiplane", "démocratie", "lord", a unifying theme is conceptual terms, often with the "dem-" prefix or related to titles.Perhaps focus on the *type* of words: specific terms, states, titles."specific 'dem' terms and titles"Let's check the length: 5 words.Does it capture it? "demoiselles", "demilitarized", "demiplane", "démocratie" all fit "specific 'dem' terms". "Lord" fits "titles".Could it be more specific by looking at the *context* from TOP_ACTIVATING_TEXTS?

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_12_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     MILLER
    0.21
     perro
    0.19
     Miller
    0.19
     cunning
    0.18
    ullen
    0.17
     Mortal
    0.17
     quiescent
    0.17
     conoz
    0.17
     curioso
    0.17
     beck
    0.17
    POSITIVE LOGITS
    SignIn
    0.21
    Squares
    0.21
    DEFGHIJKLMNOP
    0.19
    Gard
    0.19
    Incre
    0.19
    signUp
    0.19
    േഷന്‍
    0.18
    sing
    0.18
     समिट
    0.18
    Health
    0.18
    Activations Density 0.001%

    No Known Activations