© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Qwen3.5-4B
    3. 15-RES-MATRYOSHKA-65K
    4. 19144
    Prev
    Next
    INDEX
    Explanations

    The neuron is detecting negative outcomes or consequences.For example, "removed due to...", "mortality of...", "losses associated with...", "injury to...".The presence of `due`, `of`, `associated`, `for`, `by`, `when` in both `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` and indirectly `TOP_POSITIVE_LOGITS` (via `due`) suggests linking these negative states to causes or circumstances.The `TOP_POSITIVE_LOGITS` also contains some foreign characters which are often specific to certain domains like medicine or technical fields where negative outcomes are discussed.Let's re-evaluate:- `removed` + `due to`- `mortality` + `of`- `losses` + `associated with`- `injury` + `to` / `caused by` / `from`- `threatened` + `by`The core idea is that these tokens represent things that went wrong or are problematic, and they are often followed by an explanation of *why* or *how*.The Chinese characters in `TOP_POSITIVE_LOGITS` might be specific terms related to medical complications or technical failures where such negative outcomes are frequent. e.g., `éĢļè¿ĩ` (complications), `è¿ĻäºĽ` (excessive), `éϤäºĨ` (expensive/costly).Considering the common elements: negative events + their causes/context.Let's look at the words themselves:`removed`, `mortality`, `losses`, `threatened`, `dismissal`, `injury`. These are all undesirable states or actions.The tokens after suggest a consequence or reason: `due`, `associated`, `for`, `injury`, `by`.So, it's about negative events and their reasons/associations.Possible explanations:- negative outcomes and reasons- undesirable events and their causes- complications and aftermath- problems and explanations- negative states and consequencesLet's check the length and specificity."negative outcomes and reasons" (4 words) - seems good.negative outcomes and reasons

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    decoderesearch/qwen-3.5-saes/qwen-3.5-4b
    Prompts (Dashboard)
    16,384 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    æĹ¢çĦ¶
    -0.06
    ader
    -0.06
    åĩºåħ¥
    -0.05
    ropp
    -0.05
    ettes
    -0.05
    unded
    -0.05
     well
    -0.05
    anda
    -0.05
     hi
    -0.05
    åĬĽ
    -0.05
    POSITIVE LOGITS
    主è¦ģéĢļè¿ĩ
    0.07
    éϤäºĨ
    0.07
    ä¸įæĺ¯åĽłä¸º
    0.06
    éĢļè¿ĩè¿ĻäºĽ
    0.06
    ä¸įä»ħä»ħ
    0.06
     кÑĢоме
    0.06
    due
    0.06
    éϤäºĨæľī
    0.06
    ä¼ļéĢļè¿ĩ
    0.06
     помимо
    0.06
    Activations Density 0.282%

    No Known Activations