© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Qwen3.5-4B
    3. 15-RES-MATRYOSHKA-65K
    4. 18843
    Prev
    Next
    INDEX
    Explanations

    Let's analyze the provided lists to find a pattern for the neuron's behavior.* **MAX_ACTIVATING_TOKENS**: `board`, `frame`, `ys`, `regular`, `Scale`, `.io`, `)`, `the`, `interview`, `)`* **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: `instead`, `are`, `-type`, `expression`, `.`, `/`, `and`, `use`, and* **TOP_POSITIVE_LOGITS**: `1`, `u`, `|`, `ing`, `olos`, `"`, `ä¸Ģç§į`, `ad`, `ni`* **TOP_ACTIVATING_TEXTS**: Scan for common themes or specific sequences.Let's examine the connection between MAX_ACTIVATING_TOKENS and TOKENS_AFTER_MAX_ACTIVATING_TOKEN and the TOP_ACTIVATING_TEXTS.1. **board** -> `instead` in "Storyboard instead" (TOP_ACTIVATING_TEXTS example: "Storyboard instead?")2. **frame** -> `are` in "Einstein conformal frame are equivalent" (TOP_ACTIVATING_TEXTS example: "Einstein conformal frame are")3. **ys** -> `-type` in "Jeffreys-type viscoelastic film" (TOP_ACTIVATING_TEXTS example: "Jeffreys-type viscoelastic film") - This suggests `ys` might be related to the end of a word like 'viscoelastic' or 'Jeffreys'.4. **regular** -> `expression` in "regular expression" (TOP_ACTIVATING_TEXTS example: "regular expression to parse")5. **Scale** -> `.` or context around it. "Death Scale." (TOP_ACTIVATING_TEXTS example: "Death Scale.")6. **.io** -> `/` or context. "Forecast.io / Dark Sky API" (TOP_ACTIVATING_TEXTS example: "Forecast.io /")7. **interview** -> likely followed by punctuation or conjunctions. "Rachel Ashwell interview" (TOP_ACTIVATING_TEXTS example: "Rachel Ashwell interview")8. `)` -> `and` (multiple occurrences after parentheses in texts)The `TOP_POSITIVE_LOGITS` list contains `ing`, `u`, `ad`, possibly suggesting gerunds, single letters often used in placeholders, or word fragments. The `ä¸Ģç§į` is a non-Latin character, hinting at potential handling of various scripts or specific technical terms. `olos` is part of `blastospores` from a text.The most consistent pattern seems to be specific words or tokens followed by common grammatical connecters or punctuation.Many of the MAX_ACTIVATING_TOKENS are nouns that often appear before a descriptor or a conjunction.Examples:- Story**board** *instead*- Einstein conformal **frame** *are*- **regular** *expression*- Death **Scale**.- Forecast**.io** /- Rachel Ashwell **interview**The neuron seems to activate for words that are often followed by another word, punctuation, or a slash, creating compound phrases or clauses. The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` are primarily functional words and punctuation.Let's re-evaluate:The combination of MAX_ACTIVATING_TOKENS (nouns, specific identifiers like .io or Scale) and TOKENS_AFTER_MAX_ACTIVATING_TOKEN (articles, conjunctions, punctuation, descriptive words like 'expression', 'type') suggests phrases or compound terms.Let's look at specific examples:- "Storyboard instead"- "Einstein conformal frame are"- "Jeffreys-type viscoelastic film"- "regular expression"- "Death Scale."- "Forecast.io /"- "Rachel Ashwell interview"The pattern is identifying specific entities or terms that are often part of a descriptive phrase or compound noun structure.Consider `TOP_POSITIVE_LOGITS`: `ing`, `u`, `|`, `olos`, `ad`. These seem like fragments or common linguistic components.`board`, `frame`, `regular`, `Scale`, `.io`, `interview` are specific terms.The neuron seems to be identifying specific terms that are part of common technical or descriptive phrases.The phrases often involve a noun followed by another noun, adverb, or punctuation.Let's try to generalize:- `board` + `instead`- `frame` + `are`- `regular` + `expression`- `Scale` + `.`- `.io` + `/`- `interview`The pattern is identifying specific tokens that are often followed by functional tokens (prepositions, conjunctions, punctuation, articles, or descriptive nouns).Let's simplify it. What do these MAX_ACTIVATING_TOKENS collectively point towards?`board`, `frame`, `Scale`, `.io`, `interview` are all specific entities/concepts often cataloged or described further. `regular` is a descriptor.The tokens after are often separators or continuations.A good candidate for a general pattern would be specific keywords followed by common word types or punctuation.Let's look at the `TOP_POSITIVE_LOGITS` again: `ing`, `u`, `|`, `olos`, `ad`. `los` is part of `blastospores`. `ad` is a common suffix. `ing` is a gerund/participle suffix.How about "specific terms followed by continuations"?This is more than 3 words."terms followed by connectors" - 4 words. This covers board instead, frame are, regular expression, Scale."specific terms and their context" - 5 words."keywords and continuations" - 3 words."specific words and following words" - 4 words.Let's refine: The neuron doesn't just detect *any* specific word, but those that tend to start specific phrases or compound terms.The `TOP_ACTIVATING_TEXTS` strongly support this."Storyboard instead""Einstein conformal frame are""Jeffreys-type viscoelastic film""regular expression""Death Scale.""Forecast.io /""Rachel Ashwell interview"The pattern is identifying specific entities/concepts and then what follows them.What if we focus on *what kind* of things these are?`board`, `frame`, `Scale`, `.io` are often technical or measurable items. `interview` is an event.Let's try to capture the essence of what precedes the "TOKENS_AFTER_MAX_ACTIVATING_TOKEN".It seems to be specific entities or significant keywords.Consider the phrase: "entity names and identifiers" - 4 words. This fits `.io`, `Scale`, `frame`, `board`."specific keywords and their context" - 5 words.Let's look for a simpler, more direct pattern.The list `MAX_ACTIVATING_TOKENS` contains a mix of common words used technically (`board`, `frame`, `regular`), specific identifiers (`.io`), and proper nouns (`interview`).The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` are often grammatical or descriptive connectors (`instead`, `are`, `-

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    decoderesearch/qwen-3.5-saes/qwen-3.5-4b
    Prompts (Dashboard)
    16,384 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    emuk
    -0.06
    plusplus
    -0.06
    ;set
    -0.06
    诸å¦Ĥ
    -0.06
    anse
    -0.06
    ajari
    -0.06
    uteur
    -0.06
    à¸ģาย
    -0.06
    edur
    -0.06
    yanto
    -0.06
    POSITIVE LOGITS
    1
    0.06
    u
    0.05
    |
    0.05
    ing
    0.05
    ↵
    0.05
    olos
    0.05
    "
    0.05
    ä¸Ģç§į
    0.05
    ad
    0.05
    ni
    0.05
    Activations Density 0.009%

    No Known Activations