© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Qwen3.5-4B
    3. 15-RES-MATRYOSHKA-65K
    4. 18216
    Prev
    Next
    INDEX
    Explanations

    the use of specific terms related to finding or describing answers.**Explanation Derivation:**1. **MAX_ACTIVATING_TOKENS**: The overwhelming presence of "solution" (repeatedly) is the most striking pattern here. This strongly suggests the neuron is highly sensitive to this word or concept.2. **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: The tokens "of", "to", "for", "?", "\", "$" indicate common grammatical structures or punctuation that follow a word. This reinforces that "solution" is the core element being tracked. For example, "solution of...", "solution to...", "solution for...".3. **TOP_POSITIVE_LOGITS**: This list is a mix. * Chinese characters: These often appear as placeholders or noise if the model encounters text it cannot perfectly parse or if certain neurons are fine-tuned on a specific subset of data. For this specific task, they probably shouldn't be the primary driver unless they form a clear pattern *with* other English words. * English words: "ester", "unud", "erc", "feld", "honored". These are less directly related to "solution" at first glance. However, I need to consider the *context* they appear in within the activating texts.4. **TOP_ACTIVATING_TEXTS**: This is crucial for context.**solution**

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    decoderesearch/qwen-3.5-saes/qwen-3.5-4b
    Prompts (Dashboard)
    16,384 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    лик
    -0.07
    ÃŃso
    -0.07
    abeth
    -0.07
    имÑĥÑīеÑģÑĤв
    -0.07
    ejm
    -0.06
    igus
    -0.06
    ohl
    -0.06
    ðŁĺī
    -0.06
    wich
    -0.06
    ILED
    -0.06
    POSITIVE LOGITS
    å¿ħè¦ģçļĦ
    0.05
    踪影
    0.05
    人家
    0.05
    深度çļĦ
    0.05
    代表æĢ§çļĦ
    0.05
    ester
    0.04
    unud
    0.04
    erc
    0.04
    feld
    0.04
     honored
    0.04
    Activations Density 0.008%

    No Known Activations