© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Olmo-3-1025-7B
    3. 16-RES-MATRYOSHKA-65K
    4. 3191
    Prev
    Next
    INDEX
    Explanations

    It tells * "Kaiju funsen–Daigoro tai Goriasu (1972) Pink Lady no KatsudÅį Daishashin (1978) Station (film) (1981) Izakaya ChÅįji (" * "niczÄħca Mam ogromnÄħ przyjemnoÅĽÄĩ powitaÄĩ tu dzisiaj delegacjÄĻ posÅĤów i innych goÅĽci z parlamentu Nowej Zelandii oraz z Misji Nowej Zelandii do" * "×Ĺ׳×ļ ×ľ×ł×¢×¨ ×¢×ľ–פ×Ļ ×ĵר׼×ķ; ×Ĵ×Ŀ ׼×Ļ ×Ļ×ĸ×§×Ļף ׾×IJ ×Ļס×ķר" * "Winners In SFR Yugoslavia Performance by club Sources 50 godini fudbal vo makedonija 1919-1969, Football Federation of North Macedonia, 1969 http://www.rsssf" * "Goal analysis and goal operationalisation: a group intervention for the enhancement of work motivation]. Work motivation, mental well-being and compet"Let's look for patterns:* **MAX_ACTIVATING_TOKENS**: Many tokens seem non-Latin or characters with diacritics. There's "anie", "ost", "ija", "Goal". The non-Latin characters (Äģ, Äħ,Åį, ów, ×ķ) suggest a focus on non-English languages or specialized character sets.* **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: These are mostly single characters or common punctuation like 'k', 'd', 'sa', 'à¶', '(', 'Da', 'i', ';', 'analysis'. "analysis" is interesting, especially paired with "Goal".* **TOP_POSITIVE_LOGITS**: "unn", "unwanted", "h", "char", "ooting", "fier", "ish", "widespread", "oir", "esti". Some of these end in "ish" or "oir" which could relate to certain languages or word endings. "unwanted", "widespread" suggest concepts.* **TOP_ACTIVATING_TEXTS**: * The first text is in Latvian ("kundze", "gada", "oktobris"). * The second is in Lithuanian ("pasaulinio", "aplinkos"). * The third is in Slovak ("rokovania", "začalo", "Vážený"). * The fourth is gibberish/special characters. * The fifth mentions "television miniseries", "Justice", "Czech crime television series". * The sixth mentions Japanese film titles ("Goriasu"). "Daigoro" and "Goriasu" are in MAX_ACTIVATING_TOKENS. * The seventh is in Polish ("niczyca", "przyjemność", "delegacja", "posłów", "gości", "parlamentu"). * The eighth is in Hebrew. * The ninth mentions "SFR Yugoslavia", "makedonija", "Football Federation". * The tenth explicitly mentions "Goal analysis" and "work motivation".Observations:* There are many non-English languages present (Latvian, Lithuanian, Slovak, Czech, Japanese, Polish, Hebrew, Macedonian/Serbo-Croatian).* There are specific sounds or character types that appear in MAX_ACTIVATING_TOKENS that might be common in these languages (e.g., Ä, Å, Ã, ×).* The presence of "Goal" in MAX_ACTIVATING_TOKENS and "analysis" after it, along with the "Goal analysis" text, points towards a specific domain.* However, the vast majority of activating texts are in non-English languages. Some of the MAX_ACTIVATING_TOKENS like "anie", "ost", "ija" could be endings in Polish or other Slavic languages.* The non-Latin characters Äģ, Äħ, Åį, ów, ×ķ are very prominent. These often appear in Central/Eastern European languages.Let's revisit the English logits: "unwanted", "widespread". These don't strongly tie to foreign languages *specifically*, but might be concepts associated. "char" could be part of many words. "ooting" is less clear. "fier" could be part of many foreign words.Considering the diversity of languages and the specific characters/tokens in MAX_ACTIVATING_TOKENS, it's likely looking for characters or token sequences common in many non-English, particularly Central/Eastern European, languages. The "Goal analysis" might be a less dominant feature, or perhaps the neuron is more general and activated by specific character sets that *happen* to appear in these diverse texts.The tokens like Äģ, Äħ, Åį, ów, ×ķ are very specific. They are not just "foreign", but suggest certain encoding or character sets.If we look at the `TOP_POSITIVE_LOGITS`, "unn", "unwanted", "h", "char", "ooting", "fier", "ish", "widespread", "oir", "esti"."widespread" and "unwanted" are conceptual."char" is common."ish" and "oir" could hint at word endings."esti" could relate to "estonia" or word endings.The explanation needs to be concise and specific. "non-english languages" is too broad.The presence of "Goal" and "analysis" is strong, but contradicted by the overwhelming presence of diverse foreign languages in the activating texts.Let's focus on the foreign characters/tokens. They are unusual for a standard English text model.The pattern seems to be specific character sets or sequences found in non-English text.The specific tokens like Äģ, Äħ, Åį, ów, ×ķ are striking. What do they represent? They look like combinations of Latin characters with diacritics or special symbols that might appear in specific encodings or languages.Let's assume these are character sequences characteristic of specific languages.The phrase should capture this.The `TOP_POSITIVE_LOGITS` have

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    decoderesearch/olmo-3-saes/olmo-3-1025-7b
    Prompts (Dashboard)
    24,576 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    kus
    -0.11
    izzo
    -0.11
    /--
    -0.10
    Segue
    -0.10
    kah
    -0.10
    aggio
    -0.10
     seri
    -0.10
    )\<
    -0.10
     Infos
    -0.09
    observe
    -0.09
    POSITIVE LOGITS
    unn
    0.08
     unwanted
    0.08
    h
    0.07
    char
    0.07
    ooting
    0.07
     fier
    0.07
    ish
    0.07
     widespread
    0.07
    oir
    0.06
    esti
    0.06
    Activations Density 0.109%

    No Known Activations