© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 25-GEMMASCOPE-2-TRANSCODER-262K
    4. 205044
    Prev
    Next
    INDEX
    Explanations

    **Neuron Behavior Analysis: Cough as a Symptom**Based on the provided data, the neuron consistently activates around the token "cough" and "coughing" within texts describing medical conditions, particularly respiratory illnesses. The tokens following "cough" are typically punctuation or descriptors that complete a symptom list.* **MAX_ACTIVATING_TOKENS**: Predominantly "cough" and "coughing."* **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: Punctuation (`,`, `.`), conjunctions (`with`), or words that continue a description.* **TOP_ACTIVATING_TEXTS**: Features "cough" in contexts of Pneumonia, HAPE, Asthma, COVID-19 symptoms, and bronchitis.* **TOP_POSITIVE_LOGITS**: This list provides less direct insight into the "cough" symptom itself, appearing more like language identifiers or unrelated terms.The clearest pattern is the association of the neuron with the symptom of coughing, usually as part of a list of symptoms for various ailments.Therefore, the explanation should focus on this core observation.Options:1. cough2. cough symptoms3. respiratory symptoms4. coughing"cough" is the most specific and direct token."cough symptoms" is also very accurate."respiratory symptoms" is a good generalization but might lose specificity if the neuron *only* triggers on "cough" specifically.Let's re-evaluate the instruction: "Keep your explanation concise (3 to 20 words). The explanation could be a single word, or phrase, or pattern.""Your job is to explain the behavior of the neuron in a single short phrase."Given "cough" is the *maximally activating token* and appears repeatedly, and the context is clearly medical symptoms, "cough" itself is a strong candidate. However, "cough symptoms" better explains the *behavior* in context.Let's consider what kind of pattern is requested. If it's just the token, "cough". If it's the context *of* the token, "cough symptoms".The prompt says "what the neuron detects or predicts by finding patterns in lists." The pattern is the *occurrence of cough leading to symptom descriptions*.Let's check for any negative examples or other strong signals. The TOP_POSITIVE_LOGITS are varied and don't scream "cough". "heritage", "Valor", "argv", "tế", "Amit", "Cellulose", "numero", "আন্তর্জাতিক". This suggests the neuron is *highly specific* to "cough" or things very closely related, rather than a broad category represented by the logits.If the neuron strongly correlates to the token "cough" when it appears in medical symptom lists, then the most direct and specific explanation is the token itself, or a very close variant."cough" is 1 word."cough symptoms" is 2 words.The rules state "The explanation could be a single word, or phrase, or pattern." and "Just say the pattern itself".The pattern is the presence of "cough" being significant. The fact that it's followed by punctuation or continuation words suggests *it's part of a list*.Considering simplicity and specificity, and the strong presence of "cough" itself:If I must choose *one* pattern from the options:"cough" - This is the most direct observation from MAX_ACTIVATING_TOKENS."cough symptoms" - This adds context confirmed by TOP_ACTIVATING_TEXTS.Let's see if the prompt gives any hints on prioritizing token vs. contextual pattern."What the neuron detects or predicts by finding *patterns in lists*." This implies it's not just the word itself, but its context in the list.Looking at the TOKENS_AFTER_MAX_ACTIVATING_TOKEN, you often see a comma or period, indicating "cough" is an item in a list.Example: "fever, cough, chest pain"Here, 'cough' is an item. What the neuron 'detects' is the *presence of cough as a symptom*. Therefore, "cough symptoms" seems more explanatory of its *behavior* than just "cough".Final consideration: "respiratory symptoms" is good, but the data *specifically* highlights "cough". If the neuron triggers on *any* respiratory symptom, then that would be better. But it's triggered by "cough". So, "cough" or "cough symptoms" is more accurate to the provided data."cough symptoms" feels like a slightly better explanation of the neuron's *role* or *purpose* than just the token itself.cough symptoms

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_25_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     Pourquoi
    0.17
     wilde
    0.16
     Lle
    0.16
     doch
    0.15
    stackpath
    0.15
    యత్
    0.15
    ]:
    0.15
     shotgun
    0.15
     sideboard
    0.15
     analogy
    0.14
    POSITIVE LOGITS
    heritage
    0.16
    ಧ್ಯ
    0.16
    Valor
    0.16
    argv
    0.16
     tế
    0.16
    Amit
    0.16
    esity
    0.16
     Cellulose
    0.15
    numero
    0.15
    আন্তর্জাতিক
    0.15
    Activations Density 0.009%

    No Known Activations