© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Olmo-3-1125-32B
    3. 32-RES-BATCHTOPK-131K
    4. 116171
    Prev
    Next
    INDEX
    Explanations

    * stone: "sandstone" (sand + stone) * storms: (related to weather, which can involve sand/sandstorms) * castle: "sandcastle" (implied) * bag: "sandbag" (implied) * bank: "sandbank" (implied) * worm: "sandworm" (implied) * pit: "sandpit" (implied) The logits suggest things *containing* sand or related *to* sand.4. **TOP_ACTIVATING_TEXTS**: * "Silver Strand Beach and Hollywood Beach are renowned for their **sand** and" -> sand followed by 'and' * "Moonlight illuminates **sandstone** buttes..." -> sandstone (sand + stone) * "**Sand** dunes near Mongolia's Har Lake..." -> Sand followed by 'dunes' * "...neotropical **sandfly** subspecies..." -> sandfly (sand + fly) * "...beginning in March at the Albert Park Circuit and ending on 13 September at **Sandown** Raceway." -> Sandown (proper noun starting with Sand) * "...and **sandhill** crane (Grus canadensis)." -> sandhill (sand + hill) * "...along the **Sandover** Highway..." -> Sandover (proper noun starting with Sand) * "...mud or **sand**. The golfer..." -> sand followed by '.' * "Common name: **sand boas**, Kenyan **sand boa**" -> sand boa (sand + boa) * "Jerry **Sandusky** has been sent..." -> Sandusky (proper noun starting with Sand)**Pattern Identification**:The core token is "sand".The tokens following "sand" are often parts of compound words or phrases related to sand:- sand + stone (sandstone)- sand + fly (sandfly)- sand + hill (sandhill)- sand + boa (sand boas)- sand + dunes (dunes)- sand + and (and)- sand + .The top positive logits are often nouns that *could* be part of a compound word with sand: stone, castle, bag, bank, pit, worm.The activating texts reinforce this:- "sandstone"- "sandfly"- "sandhill"- "sand boas"- "Sandown" (proper noun, but starts with Sand)- "Sandover" (proper noun, but starts with Sand)- "Sandusky" (proper noun, but starts with Sand)The pattern is clear: the neuron focuses on the word "sand" and what typically follows it, forming compound words or relevant phrases/nouns.**Concise Explanation Generation**:- "sand" + [stone, fly, hill, boa, dunes, etc.]- words followed by specific things related to sand: stone, fly, hill, boa, dunes- sandstone, sandfly, sandhill, sand boas- sand and words like stone, fly, hill, boa- sand followed by specific nouns- sand + follow-up wordLet's try to be more specific. It's about the *word class* or *semantic category* of the token following "sand".- sand and stone- sand and fly- sand and hill- sand and boa- sand and dunesThe logits also give clues: stone, castle, bag, bank, pit, worm. These are all things that can be made of sand or are related to environments with sand.So, it's "sand" followed by things that modify it or are associated with it."sand and stone" -> sandstone"sand and fly" -> sandfly"sand and hill" -> sandhill"sand and boa" -> sand boaThe pattern is "sand" followed by a noun or adjective that forms a common compound.Possible explanations:- sand followed by noun compounds- sand and associated nouns- sandstone, sandfly, sandhill, sand boas (this is listing, rule says avoid)- sand and its compounds- sand + noun (stone, fly, hill, boa) - too specific- sand and related conceptsThe prompt says "find patterns that helps you explain the behavior".The *most* prominent pattern is "sand" appearing, and then specific subsequent words that form common compounds. The logits confirm these compounds are highly relevant.Let's look at the TOKENS_AFTER_MAX_ACTIVATING_TOKEN again:andstonedflyownhillover.bousThis is a bit noisy. "d" for "sanddune"? "own" for "sandown"? "over" for "sand over"? "bo" for "sand boa"? "us" for "sandus"?The TOP_POSITIVE_LOGITS are key: stone, storm, paper, castle, bag, bank, storms, worm, pit.These are nouns that commonly follow "sand", either directly or in compounds:- stone -> sandstone- castle -> sandcastle- bag -> sandbag- bank -> sandbank- worm -> sandworm- pit -> sandpitAmong the TOKENS_AFTER_MAX_ACTIVATING_TOKEN, we see 'stone', 'fly', 'hill', 'bo'. These directly map to:- stone -> sandstone- fly -> sandfly- hill -> sandhill- bo -> sand boaThe pattern is "sand" followed by words that create specific compounds.How concise can I be?"sand and compounds""sand and related nouns""sandstone, sandfly, sandhill, sandboa" (listing, avoided)"sand followed by modifying words"Let's see what the core *concept* is. It's the word "sand" plus specific subsequent tokens or concepts.The TOP_POSITIVE_LOGITS are words that *often pair with sand*."sandstone", "sandcastle", "sandbag", "sandbank", "sandworm", "sandpit".The MAX_ACTIVATING_TOKENS are just "sand".The TOKENS_AFTER_MAX_ACTIVATING_TOKEN: stone, fly, hill, bo suggest "sandstone", "sandfly", "sandhill", "sand boa".The neuron is detecting "sand" in contexts where it's part of a specific compound noun."sand followed by compounds" (3 words) - Good."sand and compound nouns" (4 words) - Good."sand plus specific noun compounds" (5 words) - Good.Let's refine. The tokens *after* sand are key."stone", "fly", "hill", "bo" seem to be the strongest signal for subsequent tokens.These map to: sandstone, sandfly, sandhill, sand boa.The neuron detects "sand" when it appears before specific other words that form common phrases or compound nouns."sand and stone, fly, hill, boa" - too long, listing.What if I just focus on the resulting compounds or concepts?The neuron identifies things like "sandstone", "sandfly", "sandhill", "sand boa", "sandcastle", "sandbag", "sandbank".The rule is "patterns".The pattern is *sand* + specific following word.Maybe "sand and specific subsequent words"? (4 words)Or "sand followed by modifiers"?

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    bcywinski/Olmo-3-32B-Base-SAE/saes_allenai_Olmo-3-1125-32B_batch_top_k/resid_post_layer_32
    Prompts (Dashboard)
    16,384 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    emit
    -0.12
    aths
    -0.11
    lices
    -0.10
    vie
    -0.10
    ROP
    -0.10
    -esque
    -0.10
    lesia
    -0.10
    iration
    -0.10
    isations
    -0.10
          ↵      ↵
    -0.10
    POSITIVE LOGITS
    stone
    0.17
    storm
    0.12
    paper
    0.12
    castle
    0.11
    bag
    0.11
    bank
    0.11
    storms
    0.10
    iego
    0.10
    worm
    0.10
    pit
    0.10
    Activations Density 0.011%

    No Known Activations