© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Qwen3.5-4B
    3. 15-RES-MATRYOSHKA-65K
    4. 21685
    Prev
    Next
    INDEX
    Explanations

    The presence of '1' followed by words like 'was', 'has', '.Column', 'and', 'as', or punctuation suggests an attempt to define or describe something, often within a structured context like code, data, or lists. The neuron might be picking up on patterns where a numeric identifier or a letter ('A', 'I') is used to preface descriptive text or structural elements.Given the tokens `1`, `A`, `I` and subsequent tokens like `.Column`, `and`, `as`, `.`, and the positive logits showing various language snippets (`porté`, `ä¹İ`, `Rencontres`, etc.), a phrase describing the neuron's focus on structured language elements or identifiers preceding descriptive text seems appropriate.Looking at `MAX_ACTIVATING_TOKENS` (e.g., `1`, `A`, `I`, `ellant`) and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` (e.g., `was`, `.Column`, `and`, `as`, `.`), a common pattern is a short token followed by a word often introducing a description or grammatical structure.Consider the example `Table1.Column1` from `TOP_ACTIVATING_TEXTS`. Here, `1` from `MAX_ACTIVATING_TOKENS` is followed by `.Column` from `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`. Another example: `Stimulus A was` - `A` is activating, followed by `was`. `script1.py` - `1` is activating, conceptually linked to file names.The `TOP_POSITIVE_LOGITS` seem like noise or from very different languages, making it harder to find a direct translation/meaning. However, `udid` and `ports` can relate to identifiers or technical terms. `Rencontres` suggests meeting/encounter.Based on `1`, `A`, `I` and `.Column`, `was`, `and`, `as`, the neuron might be detecting a pattern of an identifier or initial followed by a descriptor.Let's focus on the structure: `identifier` + `descriptor`.identifier followed by descriptor

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    decoderesearch/qwen-3.5-saes/qwen-3.5-4b
    Prompts (Dashboard)
    16,384 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    è¯ļ
    -0.06
    ãģ»ãģ©
    -0.05
     multif
    -0.05
    äºĴ为
    -0.05
     human
    -0.05
    lite
    -0.05
    ef
    -0.05
    ede
    -0.04
     prox
    -0.04
    -
    -0.04
    POSITIVE LOGITS
    porté
    0.07
    ä¹İ
    0.07
     Rencontres
    0.07
     ÐĶнем
    0.07
    udid
    0.07
     Dodatkowo
    0.06
     Potre
    0.06
     thun
    0.06
    ports
    0.06
    Åijbb
    0.06
    Activations Density 0.020%

    No Known Activations