© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B
    3. 31-GEMMASCOPE-2-RES-262K
    4. 240097
    Prev
    Next
    INDEX
    Explanations

    booking scorer visited"}, {"MAX_ACTIVATING_TOKENS": ["safely", "oil", "."]}, {"TOKENS_AFTER_MAX_ACTIVATING_TOKEN": ["engage", "Once", "Agreement"]}, {"TOP_POSITIVE_LOGITS": ["train", "year", "travel"]}, {"TOP_ACTIVATING_TEXTS": ["safely engage", "olive oil", "scorer all-time", "visited one", "training Agreement"]}]In this case, the examples show a pattern of actions or states, often followed by specific outcomes or descriptions.- `safely engage`- `olive oil`- `scorer all-time`- `visited one`- `training Agreement`The TOP_POSITIVE_LOGITS also suggest specific entities or concepts: `train`, `year`, `travel`.Considering the MAX_ACTIVATING_TOKENS and TOKENS_AFTER_MAX_ACTIVATING_TOKEN:- `to` -> `decide`- `safely` -> `engage`- `olive` -> `oil`- `.` -> `Once`- `booking` -> `of`- `scorer` -> `all`- `visited` -> `one`- `aining` -> `Agreement`The patterns `safely engage`, `olive oil`, `scorer all-time`, `visited one`, `training Agreement` seem to be key. The overall theme leans towards actions, roles, or objects followed by related descriptors or outcomes.A good concise explanation would be:"actions and states followed by descriptions" or "verbs and nouns followed by descriptors"Let's refine this to be more specific to the examples.The neuron seems to activate in contexts where an action or entity is being described or concluded."safely engage" - action + outcome/state"olive oil" - object + type"scorer all-time" - role + descriptor"visited one" - action + quantity/identifier"training Agreement" - subject + related documentIt is about specific entities or actions and what follows them.The TOP_POSITIVE_LOGITS like `teenager`, `imprinted`, `crossed`, `Disco`, `गणपति`, `वाह`, `별`, `雪` are diverse, suggesting the neuron might not be capturing a *semantic category* like "animals" or "food", but rather a *structural* or *discourse* pattern.Let's focus on the `MAX_ACTIVATING_TOKENS` and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` pairs and the `TOP_ACTIVATING_TEXTS`.`to decide``safely engage``olive oil``booking of``scorer all``visited one``aining Agreement`The pattern is often a word (or part of one) followed by a specific, highly probable word.This looks like identifying common collocations or specific phrases.Let's re-evaluate the positive logits: `teenager`, `imprinted`, `crossed`, `Disc`, `गणपति`, `वाह`, `별`, `雪`. These are very diverse, which *reinforces* the idea that the neuron is not about semantic meaning of the *types* of things, but *how* they are discussed or what follows them.The phrase "verbs and nouns followed by descriptors" captures this general idea."common word pairs and phrases" also works.Let's try to be more specific to the prompt's examples."verbs and specific subsequent words""entities and their related descriptors"How about focusing on the "finding patterns in lists" aspect.The neuron finds specific word sequences.Consider the phrase structure: [Word/Token] [Specific Following Word/Token]e.g.safely + engageolive + oilscorer + allvisited + oneaining + AgreementThis is finding specific word pairings.Let's check length constraints: 3 to 20 words."specific word sequences" (3 words) - good, but maybe too generic."identifying specific word pairings" (4 words) - better.Let's look at another example:MAX_ACTIVATING_TOKENS: `to` `safely` `olive` `booking` `scorer` `visited` `aining`TOKENS_AFTER_MAX_ACTIVATING_TOKEN: `decide` `engage` `oil` `of` `all` `one` `Agreement`This strongly suggests collocations or specific phrase components.The TOP_POSITIVE_LOGITS being diverse reinforces that it's not about classifying *what* it is, but *how* it's expressed.Phrase: "specific word pairings and sequences" (4 words).Let's check the examples from the top texts again:"safely engage""olive oil""booking of academic""scorer all-time""visited one of"The pattern is indeed picking out common, identifying sequences.Let's try: "common collocations and phrases" (4 words). This fits the observed behavior well.Let's reconsider the constraint: "You should look at the lists and find a pattern that helps you explain the behavior of the neuron."The neuron is *finding* these patterns. What *is* the pattern it finds? It finds that certain tokens are frequently followed by specific other tokens, forming common phrases or collocations."identifying common word pairings" - concise, accurate. 4 words.Let's test against rules:- Concise (3-20 words): yes, 4 words.- Single word, phrase, or pattern: yes, phrase.- About tokens following/pre

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-pt/resid_post/layer_31_width_262k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    b
    0.58
    軛
    0.52
    ك
    0.52
    A
    0.52
    Haryana
    0.50
    ृती
    0.49
    ಬ
    0.49
    सीबी
    0.48
    োস
    0.48
    墀
    0.48
    POSITIVE LOGITS
     teenager
    0.46
     वाह
    0.43
     별
    0.43
     huh
    0.43
     imprinted
    0.43
     गणपति
    0.42
     Disc
    0.42
    asibility
    0.42
     crossed
    0.41
     雪
    0.41
    Activations Density 0.001%

    No Known Activations