© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Qwen3.5-4B
    3. 15-RES-MATRYOSHKA-65K
    4. 21943
    Prev
    Next
    INDEX
    Explanations

    * This neuron seems to be strongly activated by numerical figures, quantities, and their relative positions or comparisons.* **MAX_ACTIVATING_TOKENS**: single, cents, figures, places, jump, -digit, order, digit. These are all related to numbers, magnitudes, or ordering.* **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: digits, per, a, from, growth, 8, than, rather. This confirms the numerical context. "digits" follows "-digit", "per" is often associated with monetary amounts ("cents"), "growth" is often qualified by "digit" (e.g., "double-digit growth"), and "than"/"rather" are used for comparisons.* **TOP_POSITIVE_LOGITS**: These are mostly non-Latin characters, which might point to specific encodings or a less clear pattern here.* **TOP_ACTIVATING_TEXTS**: * "high-single digits" * "0.8 cents per gallon" * "triple figures a week" * "29 places" * "three notch jump from #8 to #5" * "double-digit growth" * "10(2)-10(3)pgm(-3) order" * "upper 70s low 80s" * "one digit than with the magnetic recording medium" * "in the 380s rather than the 370s" The texts clearly show a pattern of numbers, amounts, percentages, rankings, and comparisons.**Synthesizing the pattern:**The neuron is triggered by words and phrases that represent numerical values, quantities, magnitudes, rankings, and relative positions or comparisons.**Concise explanation:**Numbers, quantities, rankings, comparisons.Let's refine this to be more like a pattern of detection.It looks for numerical values and their relations.Phrases that fit the criteria (3-20 words, specific, pattern-based):- numerical values and comparisons- quantities, rankings, digits- numerical figures and ordering- digits, amounts, and positionsLooking at the MAX_ACTIVATING_TOKENS and TOKENS_AFTER_MAX_ACTIVATING_TOKEN, the most prominent theme is numerical.- `single`, `cents`, `figures`, `places`, `jump`, `-digit`, `order`, `digit`- `digits`, `per`, `growth`, `8`, `than`, `rather`The texts reinforce this:- `high-single digits`- `0.8 cents per gallon`- `triple figures`- `29 places`- `three notch jump`- `double-digit growth`- `order` (in a positional sense)- `upper 70s low 80s`- `one digit`- `380s rather than 370s`The neuron is clearly about numbers, quantities, and comparisons involving them.Let's try to find a phrase that summarizes this well and fits the rules."numerical values and their comparisons" - 5 words. Specific."digits, amounts, rankings, comparisons" - 4 words. Specific."quantities, financial figures, and comparisons" - 5 words. Specific."numerical data and relative positions" - 5 words. Specific.The key elements are digits, quantities, and how they relate (comparing, ordering, growth).The `TOP_POSITIVE_LOGITS` are weird characters, which I

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    decoderesearch/qwen-3.5-saes/qwen-3.5-4b
    Prompts (Dashboard)
    16,384 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    upert
    -0.06
    ç»Ŀ
    -0.05
    hack
    -0.05
     Pee
    -0.05
    ä½ĻåĬĽ
    -0.05
    æĢİæł·
    -0.05
    çļĦç²¾
    -0.05
    cly
    -0.05
    odik
    -0.05
     mang
    -0.05
    POSITIVE LOGITS
    лл
    0.06
    ÑĶ
    0.06
    ész
    0.06
    asan
    0.05
    èĤ¤
    0.05
    akra
    0.05
    werks
    0.05
    egan
    0.05
    аÑģÑĭ
    0.05
    ê¶ģ
    0.05
    Activations Density 0.006%

    No Known Activations