© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 3-GEMMASCOPE-2-TRANSCODER-262K
    4. 46801
    Prev
    Next
    INDEX
    Explanations

    ".The word "Longshan" appears in the MAX_ACTIVATING_TOKENS list ("shan" is a substantial part of it), and "Longshan Temple" is explicitly mentioned in the TOP_ACTIVATING_TEXTS.The word "Bao" appears multiple times in MAX_ACTIVATING_TOKENS, and "Xiao Long Bao" (a type of dumpling) is mentioned in TOP_ACTIVATING_TEXTS.The word "jing" appears in MAX_ACTIVATING_TOKENS, and "Longjing tea" is mentioned in TOP_ACTIVATING_TEXTS.The token "Phu" appears in MAX_ACTIVATING_TOKENS, and "Long Phuoc Secondary Schools" is mentioned in TOP_ACTIVATING_TEXTS.Combining these observations:- "shan" -> "Longshan Temple"- "jing" -> "Longjing tea"- "Bao" -> "Xiao Long Bao" (dumplings)- "Phu" -> "Long Phuoc" (educational institution)The pattern seems to involve specific Chinese/Vietnamese place names or food items, often preceded by "Long" or containing a syllable that sounds like "Long".The "TOKENS_AFTER_MAX_ACTIVATING_TOKEN" show context:- "Temple" after "shan"- "(" and ")" after "Bao"- "tea" and ")." after "jing"- "oc" (which could be part of "occurs" or a suffix) after "Phu"The TOP_POSITIVE_LOGITS are very diverse and might be noise or indicate broader activation.The most consistent pattern across MAX_ACTIVATING_TOKENS and TOP_ACTIVATING_TEXTS is specific Chinese/Vietnamese proper nouns, food items, or place names, often appearing in conjunction with "Long" or similar sounds.Let's try to capture this. "Longshan", "Longjing", "Bao", "Phuoc".The neuron seems to activate for specific proper nouns, especially those found in East Asian contexts, and often associated with places or food.Consider the word "Long" itself. It's not in MAX_ACTIVATING_TOKENS but forms a prefix in several key text examples ("Longshan", "Longjing", "Long Phuoc", "Longreads").The tokens in MAX_ACTIVATING_TOKENS are parts of these longer words/names."shan" -> Longshan"jing" -> Longjing"Bao" -> Xiao Long Bao"Phu" -> Long PhuocThe neuron is identifying parts of Chinese/Vietnamese names and items, especially those starting with or containing sounds similar to "Long".Explanation needs to be 3-20 words.Possible patterns:- Chinese place names and food items- Parts of specific East Asian names or items- "Long" prefixed names and items- Names like Longshan, Longjing, BaoLet's go with something that captures the components and the context.The neuron detects specific names and food items, often prefixed with "Long". This seems to fit the criteria.Let's re-evaluate MAX_ACTIVATING_TOKENS:shan -> Longshanjing -> LongjingBao -> Xiao Long BaoPhu -> Long PhuocAnd TOKENS_AFTER_MAX_ACTIVATING_TOKEN:Temple(tea).ocAnd TOP_ACTIVATING_TEXTS:Longshan TempleLongjing (Dragon Well) teaXiao Long BaoLong Phuoc Secondary SchoolsLong readsThe neuron is sensitive to components of specific place names and food items, particularly those in East Asian contexts, often linked with the prefix "Long"."shan", "jing", "Bao", "Phu" are fragments.The explanation should focus on what these fragments *represent* in context."Long" names and food items. (5 words)Specific East Asian names and food items. (6 words)Names like Longshan, Longjing, Bao, Phu. (7 words)Parts of "Long" prefixed East Asian names/foods. (7 words)The instruction says "just say the pattern itself".Let's think about what feature these tokens share. They are all parts of specific compound nouns, mostly proper nouns or food names, from East Asian languages. Many of them have "Long" as a preceding element."Long"-prefixed East Asian names and foods. This is descriptive.The neuron identifies components of "Long" prefixed East Asian names and foods.Let's check the rules:- Concise: 3-20 words. Yes.- Single word or phrase/pattern. Yes.- About tokens following/preceding. Yes, context is considered.- About words starting with sequences. Yes.- Avoid listing all tokens. Yes.- Just say the pattern. Yes.- No "This neuron detects/predicts". Yes.- No "tokens" or "patterns". My internal thinking used "pattern", but the output phrase should avoid it.- No capitalization unless proper noun. "Long" is a proper noun part.- Specific. Yes.- Majority match. Yes.Consider the parts: shan -> Longshan (Temple). jing -> Longjing (tea). Bao -> Xiao Long Bao (dumpling). Phu -> Long Phuoc (School).The common thread is the "Long" prefix or internal part, and the specific item (Temple, Tea, Bao, School).How about focusing on the most visually distinct common element: "Long" mixed with specific item names."Long" Chinese place names and food items. (6 words)"Long" prefix names and food items. (6 words)The specific tokens are SHAN, JING, BAO, PHU. These are *fragments*. The neuron might be detecting these fragments and their common context.Let's look at the TOP_POSITIVE_LOGITS again:"سیاسیส์”,tronicitaanInterfacesiresbersamaவும்These might suggest the neuron is broadly about *named entities* or *classified entities* that can appear in different contexts/languages. But that's too broad. The MAX_ACTIVATING_TOKENS and TARGET_ACTIVATING_TEXTS are more specific.The "Long" prefix is very strong."Long" prefixed East Asian names/foods. This feels right. "East Asian" covers Chinese, Taiwanese, Vietnamese contexts implied by the names and foods.Let's consider alternative phrasing:- "Long"-prefixed proper nouns and food items. (6 words)- Named entities prefixed with "Long". (5 words)The input tokens are 'shan', 'jing', 'Bao', 'Phu'. These are *parts* of longer names/items.If the neuron fires for 'shan', it's often because it's part of 'Longshan'.If it fires for 'Bao', it's often because it's part of 'Xiao Long Bao'.So the neuron might be sensitive to these specific sub-word units that frequently occur as part of "Long"-prefixed East Asian names and food items.The explanation should be the *behavior*.Behavior: what it detects or predicts.It

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_3_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    a
    0.92
    es
    0.74
    d
    0.71
    é
    0.71
    aol
    0.68
    ed
    0.68
    s
    0.67
    in
    0.67
     in
    0.67
    e
    0.66
    POSITIVE LOGITS
    "
    0.55
     سیاسی
    0.52
    ส์
    0.51
    ”،
    0.51
    tronic
    0.50
    itaan
    0.50
    Interfaces
    0.50
    ires
    0.48
     bersama
    0.48
    வும்
    0.47
    Activations Density 0.002%

    No Known Activations