© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Olmo-3-1025-7B
    3. 16-RES-MATRYOSHKA-65K
    4. 3943
    Prev
    Next
    INDEX
    Explanations

    **Explaining Neuron Activity**I'm analyzing the provided data points to identify a recurring theme or pattern. My goal is to distill the essence of the neuron's activation into a short, descriptive phrase that adheres to your specific constraints.Here's a breakdown of the analysis:* **MAX_ACTIVATING_TOKENS:** Contains numbers like `3, 25, 10, 5, 500, 50`.* **TOKENS_AFTER_MAX_ACTIVATING_TOKEN:** Shows units and contexts related to these numbers: `years, months, days, POST, beds, verte, acute`.* **TOP_POSITIVE_LOGITS:** Includes terms like `consecutive`, `successive`, `ANY`, `localized`, suggesting continuity, thresholds, and specific contexts.* **TOP_ACTIVATING_TEXTS:** This is the most crucial for identifying the pattern. Let's look for numbers followed by time units or other quantities: * `six-month` * `2 or more years` * `âī¥3 months` * `> 25 days` * `at least 10 years` * `at least 5 POSTS` * `25 beds or more` * `at least 50 years old` * `at least 5 years of service` * `âī¥3 vertebral segments` * `at least one worker and $500,000 in annual income` * `at least 2 non-neonatal pediatric wards` * `at least 50 acute inpatient beds`The overarching pattern is the mention of quantities or durations, often above a certain threshold or minimum, followed by units like years, months, days, beds, posts, etc. The logits like 'consecutive' and 'successive' reinforce this idea of quantities or periods.Therefore, the neuron seems to be activated by numerical thresholds, durations, or minimum amounts, often in the context of time or capacity.Let's try to form a phrase:* "numbers followed by time or units" - too literal.* "minimum durations or quantities" - better.* "quantities and durations" - concise.* "thresholds and time periods" - also good.* "quantities, years, months, days" - too specific/listing.* "minimum required years or months" - too specific.* "units of time and quantity" - good.* "minimum quantities or durations" - fits the data well, capturing both explicit numbers and the 'at least'/'more' context.Considering the `TOP_POSITIVE_LOGITS` like 'consecutive' and 'successive', and the common use of numbers with time units (`years`, `months`, `days`), a phrase focusing on these temporal or quantitative measures seems appropriate. The `POST` and `beds` also suggest quantities, but time units are dominant.Let me refine based on the rules: 3-20 words, no "tokens" or "patterns", no "This neuron detects/predicts", specific.The numbers often represent minimums or counts. The context is frequently time, but also capacity or requirements."minimum quantities or durations" seems strong."thresholds and time measurements" is also good."counts and time units" is simple and covers it.Let's re-examine the `TOP_POSITIVE_LOGITS`:`consecutive`, `(>`, `successive`, `secutive`, `monic`, `ANY`, `IFO`, `localized`, `mpr`, `Foot``consecutive`, `successive` strongly point to sequences, periods, or amounts.`(`>` implies a threshold.`ANY` might suggest a broad application of these quantities.`localized` could apply to specific contexts where these are measured.minimum quantities or durations

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    decoderesearch/olmo-3-saes/olmo-3-1025-7b
    Prompts (Dashboard)
    24,576 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    odel
    -0.09
    /page
    -0.09
    rk
    -0.09
    reso
    -0.09
    iale
    -0.09
     limit
    -0.09
     geh
    -0.09
    uku
    -0.09
    antis
    -0.08
     GK
    -0.08
    POSITIVE LOGITS
     consecutive
    0.11
     (>
    0.08
     successive
    0.08
    secutive
    0.08
    monic
    0.08
     ANY
    0.07
    IFO
    0.07
     localized
    0.07
    mpr
    0.07
    Foot
    0.07
    Activations Density 0.173%

    No Known Activations