© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 14-GEMMASCOPE-2-TRANSCODER-262K
    4. 98081
    Prev
    Next
    INDEX
    Explanations

    The neuron seems to detect specific nouns or concepts, often technical or object-related, followed by common punctuation or simple words. "Bottom" and "Tiles" also appear, suggesting a range of specific items or locations. The `TOP_POSITIVE_LOGITS` list is quite varied ("crime", "gardening", "syndicate"), making it hard to pin down a semantic theme directly.However, let's focus on "ventilation". It appears in a question about an "electric cooktop with downdraft ventilation". This suggests the neuron might be looking for specific technical terms or components, possibly in a context of products or systems.Let's look at the MAX_ACTIVATING_TOKENS again: ventilation, Bottom, Tiles, of, ism, of, control, ")", .And TOP_POSITIVE_LOGITS: spree, kiln, hui, ivy, crime, gardening, plomb, kui, dt, syndicate.The prompt asks for a concise explanation (3-20 words) of what it detects or predicts by finding patterns in lists.Let's reconsider the MAX_ACTIVATING_TOKENS and see what follows them in the general text, looking for any commonality.- ventilation?- Bottom" - Ritmos variados...- Tiles (Space Shuttle):- of- ism- of- control at each point.- ")".The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` are: ?, ", 1, (, evidence, ,, content, at, is.Some of these are punctuation or common words.Let's look for patterns in `MAX_ACTIVATING_TOKENS` themselves alongside `TOP_POSITIVE_LOGITS`.`MAX_ACTIVATING_TOKENS`: ventilation, Bottom, Tiles, of, ism, of, control, ")", .`TOP_POSITIVE_LOGITS`: spree, kiln, hui, ivy, crime, gardening, plomb, kui, dt, syndicateThere isn't a clear semantic overlap between `MAX_ACTIVATING_TOKENS` (vents, bottom, tiles, control) and `TOP_POSITIVE_LOGITS` (crime, gardening, syndicate, kiln, plumbing).The most concrete evidence comes from "ventilation" which is explicitly in the `MAX_ACTIVATING_TOKENS` and appears in the text with "cooktop". "Tiles" appears with "Space Shuttle". "control" appears with "Manage by Stages" and "Manage by Exception". "Bottom" is from a song title.These seem like specific nouns or terms, often related to objects, processes, or specific places/things.Let's try to find a rule that fits.- Specific nouns: ventilation, Tiles, control, Bottom- Technical terms: ventilation, Tiles, kiln, plomb, dt- Places/Items: ventilation, Tiles, Bottom, gardening- Verbs/actions: control, spree, crime, gardeningThe instructions say: "find patterns that helps you explain the behavior of the neuron." and "explain what the neuron detects or predicts".Let's focus on what the neuron is *activating* on, which are the `MAX_ACTIVATING_TOKENS`.If we see words like "ventilation", "Tiles", "control", "Bottom", "kiln", "plomb", "gardening", "syndicate", "crime", "spree", "kui", "hui", "ivy", "dt"...These are generally specific words. Some are concrete objects or places, others are actions or concepts.The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` list (`?, ", 1, (, evidence, ,, content, at, is`) looks like very generic continuations.Let's consider the `TOP_POSITIVE_LOGITS`. These are words the neuron *likes*.The most concrete items in `MAX_ACTIVATING_TOKENS` are: ventilation, Bottom, Tiles, control.The most concrete items in `TOP_POSITIVE_LOGITS` are: crime, gardening, syndicate, kiln, plomb.The prompt states: "explain what the neuron detects or predicts by finding patterns in lists." and "look at the lists and find a pattern that helps you explain the behavior of the neuron."A rule could be: "Specific terms or nouns".Let's see if any items in `MAX_ACTIVATING_TOKENS` are particularly distinct."ventilation""Bottom""Tiles""control"What a neuron detects is a feature. This neuron seems to be detecting specific entities or concepts.Let's combine the information from `MAX_ACTIVATING_TOKENS` and `TOP_POSITIVE_LOGITS`.`MAX_ACTIVATING_TOKENS`: ventilation, Bottom, Tiles, of, ism, of, control, ")", .`TOP_POSITIVE_LOGITS`: spree, kiln, hui, ivy, crime, gardening, plomb, kui, dt, syndicateThis neuron seems to activate for specific nouns related to objects, places, or abstract concepts like "crime" or "gardening". The `MAX_ACTIVATING_TOKENS` show concrete items or processes ("ventilation", "Tiles", "control", "Bottom"). The `TOP_POSITIVE_LOGITS` extend this to actions/concepts ("spree", "crime", "gardening", "syndicate") and technical terms related to materials/industry ("kiln", "plomb", "dt").A very concise descriptor for this neural behavior could be "specific nouns and concepts".Let's check the rules again:- Concise (3 to 20 words): "Specific nouns and concepts" is 4 words.- Single word, or phrase, or pattern.- Find patterns.- Avoid listing all tokens.- No "This neuron detects/predicts", "tokens", "patterns".- Do not capitalize.What if we try to find a more descriptive phrase based on the commonality of the *types* of words?`MAX_ACTIVATING_TOKENS`: concrete nouns, processes, abstract concepts.`TOP_POSITIVE_LOGITS`: abstract concepts, specific nouns, materials, industries.Consider the words that are specific:ventilationBottomTilescontrolspreekilncrimegardeningplombsyndicateThese are specific terms, not general words like "the", "a", "is". They refer to particular things or ideas.Trying to refine "Specific nouns and concepts":- "concepts and specific terms" (4 words)- "specific terms and entities" (4 words)- "specific nouns and actions" (4 words)Let's look at the source of `MAX_ACTIVATING_TOKENS`.The example text for "ventilation" is "cooktop with downdraft ventilation?".The example text for "Tiles" is "Heat Shield Tiles (Space Shuttle):".The example text for "control" is "review and control at each point." or "Manage by Stages ... review and control".The examples strongly point towards specific items, technical terms, or processes. "Bottom", from a song title, is a specific named entity.Let's propose: "specific items and processes"If I consider

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_14_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    V
    0.45
     Toolbar
    0.44
     blamed
    0.44
     instilled
    0.43
    ولي
    0.43
    爣
    0.43
    ச்சிக்க
    0.42
    男孩
    0.42
     мама
    0.42
    ichtig
    0.42
    POSITIVE LOGITS
     spree
    0.52
     kiln
    0.50
     hui
    0.49
     ivy
    0.47
     crime
    0.45
     gardening
    0.45
     plomb
    0.45
     kui
    0.45
     dt
    0.45
     syndicate
    0.44
    Activations Density 0.000%

    No Known Activations