© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-12B
    3. 24-GEMMASCOPE-2-RES-16K
    4. 551
    Prev
    Next
    INDEX
    Explanations

    The pattern identified is that the neuron activates when certain punctuation or formatting marks (like `.`, `*`) are followed by words that indicate a continuation, a statement, or a specific subject, especially in contexts that might involve dialogue or descriptive text. The `TOP_POSITIVE_LOGITS` contain emoticons and a mix of non-English words, suggesting a connection to informal text, emotional expression, or potentially a multilingual context where these symbols are used.Given the dominance of `.` and `*` in `MAX_ACTIVATING_TOKENS` and the varied sentence starters in `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`, coupled with emoticons in `TOP_POSITIVE_LOGITS`, the neuron seems to be sensitive to the transition from punctuation/formatting to sentence beginnings.The most fitting and concise explanation, considering the rules:sentence beginnings after punctuation

    np_acts-logits-general · gemini-2.5-flash-lite

    The neuron detects emphatic or evaluative discourse markers—those bold opinion or transition phrases (e.g. “No wonder,” “Let me be clear,” “Absolutely,” “It’s time to…”) that signal strong author emphasis.

    oai_token-act-pair · o4-miniTriggered by @jyhe0408
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-12b-pt/resid_post/layer_24_width_16k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    これにより
    0.54
     بنابراین
    0.47
     Additionally
    0.46
    Additionally
    0.44
    ↵
    0.43
     ซึ่ง
    0.43
     நிலையத்தில்
    0.41
    dengan
    0.41
    そのため
    0.41
    రాజు
    0.41
    POSITIVE LOGITS
    oooo
    0.55
     ;)
    0.55
     wszyscy
    0.54
     niemand
    0.54
     ყველა
    0.53
     zumindest
    0.53
     choć
    0.52
     EVERY
    0.51
     😉
    0.51
     większość
    0.51
    Activations Density 0.161%

    No Known Activations