© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 10-GEMMASCOPE-2-TRANSCODER-262K
    4. 24396
    Prev
    Next
    INDEX
    Explanations

    reasons", "your reason", "your Reason""The weather is nice today" ("今天天气不错啊")"inference"Russian text (ending in "лирование")There seems to be a mix of English words like "Reason", "explanations", "inference", and non-English words/phrases.The phrase "The weather is nice today" is very specific.The presence of "Reason" and "explanations" suggests a need for justification or information."inference" points to model predictions."ли" and "不错" suggest multilingual capabilities.Let's try to find a unifying theme."Reason" and "explanations" are about providing causes or details.The Chinese phrase "今天天气不错啊" also has a sense of observation/description."inference" is about prediction.The prompt asks what the neuron detects *or predicts by finding patterns in lists*.The neuron seems associated with generating text that:1. Explains a reason or provides information ("Reason", "explanations").2. Describes observations or simple statements ("The weather is nice today").3. Relates to the model's internal workings ("inference").4. Handles multiple languages ("ли", "不错", "啊").Considering the `TOP_POSITIVE_LOGITS`, they are mostly abstract, complex English words.explanations and reasons

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_10_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    т
    0.58
     т
    0.51
    Shelf
    0.45
    섭
    0.45
     t
    0.44
    EOS
    0.44
    ície
    0.43
    Ger
    0.43
    Circular
    0.41
     Sponsors
    0.41
    POSITIVE LOGITS
     emblems
    0.62
     argomento
    0.54
     operas
    0.53
     explanations
    0.53
     abortions
    0.53
     engravings
    0.53
     carbonate
    0.52
     arrhythmias
    0.52
     vasomotor
    0.52
     lemmas
    0.50
    Activations Density 0.000%

    No Known Activations