© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B
    3. 31-GEMMASCOPE-2-RES-262K
    4. 177894
    Prev
    Next
    INDEX
    Explanations

    places and contexts

    np_acts-logits-general · gemini-2.5-flash-lite

    categories and difficulty levels (M)2. directions or types (order type, altcoin)3. supervised or automated processes (supervised work, automated facilities)4. related to specific tokens like "altcoin", "facilities", "supervised", "category", "direction", "M".The strongest recurring pattern seems to be about specifying types or classifications of things, especially in technical or administrative contexts.Let's try to find a unifying theme."category M", "direction", "alt coin", "supervised", "automated facilities".This neuron seems to be distinguishing between different categories, types, or states.The presence of "M" (often a difficulty level), "category", "direction", "alt" (as in altcoin), "supervised", "automated", and "facilities" suggests it's classifying or categorizing information.The most specific and recurring elements point towards classifying entities, states, or types.Keywords: category, M, direction, alt, coin, supervised, automated, facilities.The prompt says "finding patterns in lists".`MAX_ACTIVATING_TOKENS` includes: `category`, `M`, `direction`, `alt`, `one`, `of`, `automated`.`TOKENS_AFTER_MAX_ACTIVATING_TOKEN` includes: `_`, `)`, `=`, `coin`, `supervised`, `of`, `where`, `facilities`.Looking at combinations:`category` + `_` -> `category_``direction` + `_` -> `direction_``M` followed by `)` or `H` in texts`alt` + `coin` -> altcoin`supervised` + `work` -> supervised work`automated` + `facilities` -> automated facilitiesThe neuron seems to be detecting labels that categorize or specify types of items, states, or processes.Considering the Top Positive logits: wildflower, deductible, slice, какой (Russian for what/which). These don't strongly align with the activation lists, suggesting the logits might be broader. The activation lists are more concrete.Let's focus on the activations.The neuron activates on things like specific categories (indicated by "category_"), difficulty levels ("M"), directions ("direction_"), types ("altcoin"), states ("supervised", "automated"), and objects ("facilities").A concise phrase could capture this idea of classification or type identification."classifies types or states" (5 words)"identifies categories and types" (4 words)"classifies technical or state terms" (5 words)"identifies specific categories or types" (5 words)Let's re-evaluate `MAX_ACTIVATING_TOKENS`: `, category M direction alt of one of automated``TOKENS_AFTER_MAX_ACTIVATING_TOKEN`: `forced not _ ) = coin supervised of where facilities`The `TOP_ACTIVATING_TEXTS` show examples that are very specific:- `category_`- `M-H`- `direction_to_count`- `altcoins`- `supervised work`- `automated facilities`The pattern is about assigning labels or classifications to entities or processes. It's detecting terms that define type, status, or category."identifies type or status labels" (5 words)"classifies categories and states" (4 words)Given the prompt is to "explain what the neuron detects or predicts by finding patterns in lists", and the lists strongly point to classifying terms, "classifies categories and states" seems a good fit. It covers "category", "M" (state of difficulty), "direction" (type), "altcoin" (type), "supervised" (state), "automated" (state), "facilities" (type/category of entity).Let's check the constraints: 3 to 20 words. "classifies categories and states" is 4 words. No intro phrases. No mention of tokens. Not capitalized unless proper noun. Specific. A majority match.The logits are a bit divergent, but the activation lists are quite consistent.The core idea is *classification*.Consider "classifies terms by type or state". (6 words)"detects classification labels" (3 words)"classifies categories and states" is concise and covers the main elements well.Let's try to make it even more specific if possible.The activation `alt` followed by `coin` is specific. `supervised` and `automated` are states. `category` and `M` are classifications. `direction` is a type. `facilities` are entities."labels things by type, state, or category" (7 words)"classifies types, states, and categories" (5 words)This is quite good. Let's stick with "classifies types, states, and categories".Final check. Explanation is "classifies types, states, and categories".- Concise (5 words, between 3-20). Yes.- Captures what neuron detects/predicts, finding patterns. Yes.- About tokens following/preceding, words starting, patterns. Yes.- Avoids listing all tokens. Yes.- No "words related to", "concepts related to". Yes.- No "This neuron detects/predicts". Yes

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-pt/resid_post/layer_31_width_262k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     Motto
    0.57
     Curran
    0.57
    梅里
    0.53
    翟
    0.52
     Munt
    0.52
    郓
    0.52
     Tra
    0.51
     Larson
    0.49
     Gö
    0.49
     Noise
    0.48
    POSITIVE LOGITS
    ся
    0.46
    ח
    0.45
     deductible
    0.45
     rice
    0.44
     questi
    0.44
     strawberries
    0.43
     રી
    0.43
     familiari
    0.43
    🫐
    0.43
     accur
    0.42
    Activations Density 0.000%

    No Known Activations