© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-4B-IT
    3. 3-GEMMASCOPE-2-TRANSCODER-262K
    4. 8740
    Prev
    Next
    INDEX
    Explanations

    address on file, deposit, on file with the card issuer, rye, on silent, deposit at the central bank, deposit at the Fed, save. * The texts mention "address on **file**", "on **file**. Information regarding", "address on **file**", "on **rye**", "on **silent**", "deposits that they are required to hold in reserve (either as cash in their vault or on **deposit**", "deposits that they are required to hold in reserve – either as vault cash or on **deposit** at the Fed", "save".**Analysis:*** **`MAX_ACTIVATING_TOKENS`**: "file" appears multiple times, "deposit" appears twice, "save" also appears. The presence of "rye" is interesting, and "silent", "Mount" are other nouns.* **`TOKENS_AFTER_MAX_ACTIVATING_TOKEN`**: Punctuation and short connecting words like "is", ".", "with", ",", "at", "(". This suggests the activated tokens are often followed by descriptions, clauses, or sentence endings.* **`TOP_ACTIVATING_TEXTS`**: The pattern "on file" is very strong ("address on file"). Other phrases are related to financial "deposits" and saving ("save"). "rye" and "silent" also appear in specific contexts.* **`TOP_POSITIVE_LOGITS`**: Mentions of "無料で" (free of charge) and general non-English characters. This might imply the neuron isn't *solely* about English file/deposit/save, or that the underlying concept can be expressed in different languages.**Pattern Identification:**The most consistent pattern across `MAX_ACTIVATING_TOKENS` and `TOP_ACTIVATING_TEXTS` is related to information storage ("file"), financial transactions ("deposit", "save"), and specific items/states ("rye", "silent"). The phrase "on file" is highly recurring.The explanation should capture this. "on file" is too literal and short. "information on file" is better, but still a bit verbose. "deposit" and "save" are key actions.Let's consider the combination of common words. "file", "deposit", "save" are strong candidates. The context often suggests a state of being or a location ("on file", "on deposit").How about focusing on the *type* of information?"address on file" strongly suggests identity or record-keeping."deposit" and "save" suggest financial or data management."rye" is an outlier but appears in a specific food context."silent" appears in the context of reducing distractions.The core seems to be about records, financial actions, or specific states."file", "deposit", "save".Let's try to combine: "file, deposit, save" - 3 words."file, deposits, save" - 3 words."on file, deposits, save" - 4 words.The `TOP_POSITIVE_LOGITS` list has non-English elements, but the prompt asks to avoid specific token listing. Let's stick to the english patterns.Considering the prompt: "finding patterns in lists", "concise explanation (3 to 20 words)", "could be about words starting with a sequence".The most prominent recurring sequence seems to be "on file". Others are discrete terms like "deposit" and "save".Let's think about what *kind* of things are activated:- address on file- data on file- deposit (bank)- save (data, money)- rye (food item)- silent (state)The commonality is not super strong across *all* items. However, "file", "deposit", "save" are repeated or heavily implied. The context of "file" and "deposit" strongly links to record-keeping, financial data, or storage.Could the pattern be related to *storing* information or items?- Address on file (stored address)- Deposit (stored money)- Save (store data/money)- file (a record/storage)- silent (a state of not broadcasting/storing active output?)Let's re-evaluate the prompt's examples for guidance. "unique words", "foreign words" are too general. Specificity is key.The literal words "file", "deposit", "save" are very common in the lists.The `TOP_POSITIVE_LOGITS` might indicate the neuron can generalize across languages or specific character sets.Let's consider "file", "deposit", "save" as the core.The phrase "on file" is key.Let's try to describe the *function* indicated by these terms.- Recording information: "address on file"- Financial transactions: "deposit", "save"- Keeping something: "file" (as in a file folder), "deposit" (as in bank deposit)Maybe it's about *holding* or *storing* something."holding files, deposits, saves" - 4 words. This is a bit awkward.How about focusing on the most frequent/prominent POSITIVE LOGITS and MAX ACTIVATING TOKENS?MAX_ACTIVATING_TOKENS: file, file, file, rye, Mount, silent, deposit, deposit, message, saveTOP_POSITIVE_LOGITS: 㐄, kem, 無料, 基本的に, ケオ, दारा, Տ, , 鲫, goreThe logits contain "無料で" which means "free of charge". This hints at financial or transactional aspects, which aligns with "deposit" and "save".The word "file" is repeated 3 times in `MAX_ACTIVATING_TOKENS` and is central to `TOP_ACTIVATING_TEXTS` ("address on file").A good explanation might be "file, deposit, save". It's 3 words and captures the most common themes from the lists. The "rye", "Mount", "silent" are less frequent but might be part of a broader context the neuron handles. The non-English logits suggest a broader linguistic capability for these concepts."file, deposit, save" seems to be the most direct interpretation of common elements.It fits the length requirement (3 words).It's specific.It avoids forbidden phrases.Let's consider if "on file" is a crucial pattern to convey."address on file""on file with the card issuer"Maybe "information on file". But that's not truly a pattern from the tokens themselves.The prompt says "The explanation could be about words starting with a sequence.""file" is a strong sequence starter. "deposit" and "save" are strong individual tokens.Alternative: "file, deposits, saving" -> 3 words.Alternative: "file, deposit, free" (referencing logits) -> 3 words.Let's stick to the tokens that appear most frequently as the strongest signal.file (3x in MAX, prominent in TEXTS)deposit (2x in MAX, prominent in TEXTS)save (1x in MAX, prominent in TEXTS)The combination captures the core. "message" is also in MAX but less emphasized elsewhere."file, deposit, save" seems robust.Let's check rules again:- Concise (3-

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-4b-it/transcoder_all/layer_3_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    ங்கிணை
    1.48
    oretically
    1.36
     trình
    1.35
    та
    1.32
    ലേ
    1.32
    ات
    1.31
    ه
    1.27
    هَا
    1.27
    ли
    1.24
    на
    1.21
    POSITIVE LOGITS
    㐄
    1.38
    kem
    1.33
    無料で
    1.22
    基本的に
    1.21
    কেও
    1.20
    दारा
    1.20
    Տ
    1.20
    
    1.20
    鲫
    1.18
     gore
    1.17
    Activations Density 0.035%

    No Known Activations