© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Olmo-3-1125-32B
    3. 32-RES-BATCHTOPK-131K
    4. 114011
    Prev
    Next
    INDEX
    Explanations

    The user wants me to identify a pattern from the provided lists and explain the neuron's behavior in 3-20 words.Here's the analysis:* **MAX_ACTIVATING_TOKENS**: Cl, Json, Jo, Http, Mut, App, Http, Ser, Te, Vari* **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: r, Data, jo, Get, ant, Kit, Host, ilog, ase, ance* **TOP_POSITIVE_LOGITS**: 象, pects, izophren, sembler, enticated, orious, Ĺı, owing, oretical, обав* **TOP_ACTIVATING_TEXTS**: Examples include C# code for `GenerationData` (starting with `Cl`), JSON handling (`JsonData`), names like `John`, `Jojo`, and `John`, HTTP requests (`HttpGet`, `HttpResponse`, `HttpClient`, `HttpHost`), `mutant/resolve`, `AppKit`, `Xamarin.Forms`, `Serilog`, `Newtonsoft.Json`, and "Variance".**Observation:**The `MAX_ACTIVATING_TOKENS` list contains initial parts of various types of technical terms, code elements, and names.The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` list shows common completions for these initial tokens.* Cl -> Clr (in `ClrDataAddress`)* Json -> JsonData* Jo -> Jojo (name)* Http -> HttpGet, HttpHost* Mut -> mutant* App -> AppKit* Ser -> Serilog* Te -> (less clear, but 'Tease' shows up)* Vari -> VarianceThe `TOP_POSITIVE_LOGITS` are a mix of specific words and what look like non-English characters or garbled text, which can sometimes indicate handling of various encodings or less common terms.The `TOP_ACTIVATING_TEXTS` confirm the presence of programming contexts (C#, JSON, HTTP, `mutant`, `Serilog`, `AppKit`), names (`John`, `Jojo`), and specific concepts (`Variance`).The core pattern seems to be identifying the *beginning* of technical terms, code constructs, or specific words that are then completed. The specific tokens like `Cl`, `Json`, `Http`, `Mut`, `App`, `Ser`, `Vari` are common prefixes or short identifiers in programming and data contexts.A good explanation should capture this idea of *initial parts leading to specific completions or concepts*.Let's try to form a phrase:1. *Recognizing initial parts of technical terms and names.* (10 words) - Good, but maybe a bit dry.2. *Identifying prefixes for code, data, and names.* (8 words) - Concise and covers the key areas.3. *Starts of programming terms and proper nouns.* (7 words) - Also good.4. *Detecting code component beginnings and names.* (6 words) - Short and captures it.5. *Code prefixes, names, and technical terms beginnings.* (7 words) - Captures variety.Looking at the `TOP_POSITIVE_LOGITS`, words like `pected`, `sembler`, `enticated`, `orious`, `oretical` suggest more complex semantic roles, but the *activations* and *following tokens* strongly point to structural

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    bcywinski/Olmo-3-32B-Base-SAE/saes_allenai_Olmo-3-1125-32B_batch_top_k/resid_post_layer_32
    Prompts (Dashboard)
    16,384 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    greSQL
    -0.33
    jamin
    -0.29
    efeller
    -0.29
    coln
    -0.28
    achusetts
    -0.28
    etheless
    -0.27
    adays
    -0.26
    gerald
    -0.24
    gomery
    -0.22
    xiety
    -0.22
    POSITIVE LOGITS
    象
    0.20
    pects
    0.16
    izophren
    0.15
    sembler
    0.15
    enticated
    0.15
    orious
    0.14
    Ĺı
    0.14
    owing
    0.14
    oretical
    0.13
    обав
    0.13
    Activations Density 2.251%

    No Known Activations