© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-12B
    3. 24-GEMMASCOPE-2-RES-16K
    4. 14679
    Prev
    Next
    INDEX
    Explanations

    These are predominantly words ending in "-ity", "-ion", "-ment", or "-ance" that describe processes, states, or abstract concepts, often related to technology, data, or analysis.Let's look at `MAX_ACTIVATING_TOKENS` again:`productivity`, `visually`, `ization`, `collection`, `progress`, `overflow`, `operation`, `manipulation`, `aspect`Observations:- Many end in "-ity" (productivity), "-ion" (ization, collection, operation, manipulation), "-ect" (aspect).- They describe abstract concepts, processes, or attributes.Now consider `TOP_POSITIVE_LOGITS`:`a`, `crunchy`, `crappy`, `情况下`, `hapless`, `flimsy`, `Roblox`, `ется`, `innocuous`, `bạn`These are very diverse and don't immediately form a strong pattern with the `MAX_ACTIVATING_TOKENS`. Some are general descriptors, some are foreign words/endings, `Roblox` is a specific entity. This list might be less helpful for a specific pattern.Let's re-examine `MAX_ACTIVATING_TOKENS` and `TOP_ACTIVATING_TEXTS`:`TOP_ACTIVATING_TEXTS` contains phrases like:- "increased efficiency and productivity"- "identifying patterns and themes visually"- "data collection"- "progress"- "buffer overflow"- "manipulating data"- "fashion aspect"Common themes emerging from `MAX_ACTIVATING_TOKENS` and `TOP_ACTIVATING_TEXTS`:- Concepts related to "doing" or "being" in a certain state.- Often related to abstract concepts, processes, or qualities.- Many end in suffixes that indicate a state or action.Let's look more closely at the suffixes:- `productivity` (-ity)- `visually` (-ly, but denotes an aspect)- `ization` (-ization, a process)- `collection` (-ion, a process/state)- `progress` (-ess, a state/process)- `overflow` (state/process)- `operation` (-ion, a process/state)- `manipulation` (-ation, a process)- `aspect` (-ect, a state/quality)The explanation should capture this focus on abstract states, processes, or qualities.The common structure is words that denote a process, state, or abstract concept, often ending in suffixes like '-ity', '-ion', '-ment', '-ance', '-ess', '-ect'.Let's try to find a concise phrase."Words ending in -ity, -ion, -ment" - too specific on suffixes, and not all fit."Abstract concepts and processes" - good, but maybe too general."States, processes, and abstract qualities" - getting closer.Looking again at `MAX_ACTIVATING_TOKENS`:`productivity``visually``ization``collection``progress``overflow``operation``manipulation``aspect`These are all nouns or adverbs functioning nominally here. They all represent abstract entities rather than concrete objects. They are states, properties, or processes.How about focusing on the common *function* rather than form? They describe *what something is* or *what is happening*.Consider the prompt's requirement: "finding patterns in lists".The pattern is that these are abstract nouns or adverbs describing states, processes, or qualities.Let's try to shorten the description."abstract states and processes" - 4 words. This seems good."abstract qualities and processes" - 4 words. Also good."states, processes, and actions" - 4 words."abstract nouns for states and processes" - 6 words.Let's check the `TOP_POSITIVE_LOGITS` again.`a``crunchy` (adjective)`crappy` (adjective)`情况下` (Chinese for "in the situation", phrase)`hapless` (adjective)`flimsy` (adjective)`Roblox` (proper noun, entity)`ется` (Russian verb ending)`innocuous` (adjective)`bạn` (Vietnamese pronoun "you")This list is still very noisy and doesn't strongly contradict the idea of abstract concepts/processes. `Roblox` is an entity, but most others are descriptive.The core of the `MAX_ACTIVATING_TOKENS` is abstract nouns denoting states, operations, or qualities.Let's try to be more specific to the examples.`productivity`, `progress`, `collection`, `operation`, `manipulation`, `overflow` are all about processes or states of being/doing.`visually`, `aspect` are about qualities or how something appears.The pattern is words describing abstract states, operations, or qualities.Let's aim for conciseness (3-20 words)."abstract states, operations, and qualities" - 5 words. This seems to cover most."abstract qualities, states, and processes" - 5 words. This is very similar.Let's try to simplify. What's the *most* common thread?`productivity` (state/result)`visually` (quality/how)`ization` (process)`collection` (process/state)`progress` (state/process)`overflow` (state/process)`operation` (process/state)`manipulation` (process)`aspect` (quality/feature)They are all abstract representations of *how* something works, *what it is*, or *what it's doing*.Consider the explanation "concepts about states and processes". This is 5 words.It fits `productivity`, `progress`, `collection`, `operation`, `manipulation`, `overflow` well.`visually` and `aspect` fit under "states" or could be seen as qualities related to states/processes.Let's check if I can make it more specific without being too narrow.The `TOP_ACTIVATING_TEXTS` often talk about:- AI's role (efficiency, productivity, patterns, data collection, algorithms)- System behavior (overflow, operation, manipulation)- Qualities/features (visually, aspect)"Abstract states and processes in systems" - 6 words. This feels quite good."Abstract qualities, states, and processes" - 5 words.I will prioritize the most common *type* of word and the *topic*.The words are abstract nouns/adverbs. They describe states, processes, or qualities.Let's think about the core meaning:`productivity`: state of being productive`visually`: how something appears`ization`: forming something`collection`: gathering things`progress`: forward movement`overflow`: excess`operation`: functioning`manipulation`: altering`aspect`: featureThese are all abstract conceptualizations. The commonality is not just that they are abstract,

    np_acts-logits-general · gemini-2.5-flash-lite

    The neuron is picking out numerical tokens—dates, years, quantities, and other numbers—in the text.

    oai_token-act-pair · o4-miniTriggered by @jyhe0408
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-12b-pt/resid_post/layer_24_width_16k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     and
    0.70
     or
    0.57
    .
    0.54
     because
    0.53
    .”
    0.53
     или
    0.51
    .’
    0.50
    .*
    0.50
    ."
    0.47
    *.
    0.47
    POSITIVE LOGITS
     a
    0.49
     the
    0.46
     yksi
    0.42
     an
    0.41
     vollständ
    0.40
     meticulous
    0.39
     OpenCV
    0.39
     Chennai
    0.38
     TensorFlow
    0.38
    айл
    0.38
    Activations Density 0.002%

    No Known Activations