© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 37-GEMMASCOPE-2-TRANSCODER-262K
    4. 126596
    Prev
    Next
    INDEX
    Explanations

    This neuron activates when "因" is followed by specific characters, often forming specific Chinese phrases.Let's break this down:- **MAX_ACTIVATING_TOKENS:** The dominant token is "因" (yīn), appearing 7 times. Other tokens are "引" (yǐn) and "应" (yìng). These are Chinese characters.- **TOKENS_AFTER_MAX_ACTIVATING_TOKEN:** - Following "因": "城" (chéng - city), "果" (guǒ - fruit/result). - Following "引": "人" (rén - person/people), "门" (mén - door/gate). - Following "应": "门" (mén - door/gate), "届" (jiè - session/period).- **TOP_POSITIVE_LOGITS:** This list contains words from various languages (Russian, Arabic, English fragments, Japanese). This suggests the neuron might be sensitive to non-Chinese text or very general concepts. However, the prompt emphasizes finding patterns in *lists*. The Chinese characters in MAX_ACTIVATING_TOKENS are much more prominent.- **TOP_ACTIVATING_TEXTS:** - "因城施策" (yīn chéng shī cè) - "city-specific policies". This example explicitly shows "因" followed by "城". - "因果" (yīn guǒ) - "karma/cause and effect". This shows "因" followed by "果". - "...引人入胜的情节..." (yǐn rén rù shèng de qíngjié) - "...engaging plot...". This shows "引" followed by "人". - "...引人入胜的对话..." (yǐn rén rù shèng de duìhuà) - "...engaging conversations...". This shows "引" followed by "人". - "...应门铃..."

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_37_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    意的
    0.49
    fieldType
    0.46
    意
    0.45
    alty
    0.40
    インスト
    0.38
     ഒഴ
    0.38
    喜欢的
    0.37
    のこ
    0.37
     Alkaline
    0.37
    tics
    0.36
    POSITIVE LOGITS
    ДУ
    0.38
    ادو
    0.37
     Hap
    0.37
     تضم
    0.37
     spann
    0.37
    дума
    0.36
     Happ
    0.36
    Camel
    0.36
    椋
    0.35
    विड
    0.34
    Activations Density 0.000%

    No Known Activations