© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Google DeepMind · Exploring Gemma 2 with Gemma Scope
    3. Gemma-2-9B-IT
    4. Residual Stream - 131k
    5. 9-GEMMASCOPE-RES-131K
    6. 15548
    Prev
    Next
    INDEX
    Explanations

    terms related to the concept of organic carbon and nitrogen within soil studies

    oai_token-act-pair · gpt-4o-miniTriggered by @bot

    donning, don't, or DON*Self-correction*: The instruction says "Do not capitalize the first letter unless it is a proper noun." "donning, don't, or DON" might be interpreted as listing words. I need to capture the *behavior*. The behavior seems to be identifying "don" and its variations, often followed by specific endings or appearing as abbreviations/names.Let's look at the *endings* again:- don .- don nees (donnees)- don ning (donning)- don ,- don )- don ' (don't)- don ka- don tidae (Notodontidae)- don Cu (Cuervo)The most common pattern isn't just "don" but specific common words or abbreviations/names associated with "don".Looking at TOP_POSITIVE_LOGITS, it lists "don", "DON", "Don", "Dons", "Donner", "Donnelly", "Дон". These are all variations of "Don" or words starting with it.The TOP_ACTIVATING_TEXTS show "DOC/DON", "donning", "don't know", "Don Armage", "Don Cuervo", "Notodontidae" (where DON is part of the name. The neuron might be detecting instances of 'don' or 'DON' in different contexts.The pattern seems to be identifying 'Don' and its related forms.Possible explanations:- Don variations- Don, DON, don't, D.O.N.- words starting with don"words starting with don" is too general."Don variations" seems reasonable."Don, DON, don't, D.O.N." is too specific a list.Let's re-examine the specific endings observed:- DON (dissolved organic nitrogen) -> followed by , or ratio or blank.- don (french 'donnees') -> followed by Courantes) or ;- donning -> followed by wild man John- don't -> followed by know- Notodontidae -> preceded by DON, or the word 'Don' itself.- Don Armage / Don Cuervo -> Proper nouns.The specific pattern seems to be distinguishing "DON" (the abbreviation) from other forms like "donning" or "don't", or specific names.The highest logits are "don", "DON", "Don". The top activating texts show "DOC/DON", "donning", "don't know", "Don Armage", "Don Cuervo", "Notodontidae".The simplest, most encompassing pattern that captures the essence of "don" across these different forms, including abbreviations like DON, common words like donning/don't, and proper names, is likely identifying the occurrence of "Don" or "don" in various forms.Let's consider the possibility that the neuron detects the substring "don".- MAX_ACTIVATING_TOKENS is all "don".- TOKENS_AFTER_MAX_ACTIVATING_TOKEN gives endings. Some are punctuation, some are word parts.- TOP_POSITIVE_LOGITS is "don", "DON", "Don", "Dons", "Donner", "Donnelly", "Дон".- TOP_ACTIVATING_TEXTS: "DOC/DON", "donnees", "donning", "don't", "Don", "donka", "Notodontidae".The neuron is clearly focused on the 'don' sequence. It appears in the prefix of many words or as an abbreviation.The endings can be varied, but the core is 'don'.Perhaps it's about specific types of 'don' words or phrases."Don" and "DON" seem prominent."donning" and "don't" are common words."Don Armage", "Don Cuervo", "Donner", "Donnelly" are names."Notodontidae" is a family name.The most consistent signal is the presence of the string "don", whether as an abbreviation "DON", part of a common word like "donning" or "don't", or as the start of a proper noun like "Don".A concise phrase:- Don variants- Don and DON- words starting with don (too general)- Don, DON abbreviationConsider "Don" as a base. It can be a title, part of a proper name, an abbreviation (DON).The neuron sees "don" followed by common endings or it sees "DON" as an abbreviation.The prompt emphasizes finding patterns in *lists*.MAX_ACTIVATING_TOKENS: DON, don, Don (capitalization variations)TOKENS_AFTER_MAX_ACTIVATING_TOKEN: ., nees, ning, ,, ), ', ka, tidae, Cu (varied endings)TOP_POSITIVE_LOGITS: don, DON, Don, Dons, Donner, Donnelly, Дон (more variations, including international)TOP_ACTIVATING_TEXTS: DOC/DON, donnees, donning, don't, Don X, Notodontidae (contextual examples)The core part is "don". The variations and endings are what differentiate its usage.The neuron might be trying to identify 'don' as a prefix or abbreviation. Phrase: "Don and DON" captures two major forms observed. Phrase: "Don prefix or DON abbreviation" is more descriptive but longer. Phrase: "Don prefix, DON abbreviation" Phrase: "Don and DON contexts"Let's try to be

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Comparing With GEMMA-2-9B-IT @ 9-gemmascope-res-131k
    Configuration
    google/gemma-scope-9b-it-res/layer_9/width_131k/average_l0_121
    Prompts (Dashboard)
    24,576 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    Features
    131,072
    Data Type
    float32
    Hook Name
    blocks.9.hook_resid_post
    Hook Layer
    9
    Architecture
    jumprelu
    Context Size
    1,024
    Dataset
    monology/pile-uncopyrighted
    Activation Function
    relu
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     CreateTagHelper
    -0.64
    fieldType
    -0.57
    .*")]
    -0.56
    IsMutable
    -0.55
    uxxxx
    -0.52
     Chwiliwch
    -0.50
     externi
    -0.50
    jší
    -0.47
     дописавши
    -0.46
    setViewportView
    -0.44
    POSITIVE LOGITS
    don
    2.06
    DON
    1.73
    Don
    1.72
     Don
    1.64
     DON
    1.52
     don
    1.28
     Dons
    1.09
     Donner
    1.06
     Donnelly
    1.05
     Дон
    1.05
    Activations Density 0.013%

    No Known Activations