© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 32-GEMMASCOPE-2-TRANSCODER-262K
    4. 236290
    Prev
    Next
    INDEX
    Explanations

    * `temporary_acting`: A boolean or flag indicating whether the head is permanently assigned to the department or is acting temporarily. (e.g., TRUE/FALSE, 1/0, 'Yes'/'- It wasn't always clear *what* `git checkout` was doing based on the arguments provided. Was it switching branches, restoring a file, or creating a new branch?This neuron detects patterns related to descriptions of states or conditions, often presented as choices or classifications.To be more specific:- MAX_ACTIVATING_TOKENS: department, branches, местности, singular, equal, in, institution, cats, id, points- TOKENS_AFTER_MAX_ACTIVATING_TOKEN: or, ,, ,, or, or, or, or, ,, or, orLooking at the MAX_ACTIVATING_TOKENS and TOKENS_AFTER_MAX_ACTIVATING_TOKEN together, there isn't a strong, consistent structural pattern like "word X followed by punctuation Y" that covers most cases. The follow-up tokens are a mix of "or" and ",".Let's examine **TOP_ACTIVATING_TEXTS** for semantic clues.- `temporary_acting`: A boolean or flag indicating whether the head is permanently assigned to the **department** or is acting temporarily.- Was it switching **branches**, restoring a file, or creating a new branch?- живешь ли ты в сельской **местности**, в городе или рядом с природой. (Russian: "do you live in a rural area, in the city or near nature.")- No Plural Markers (usually): You often rely on context to understand if something is **singular** or plural.- way to check for string **equal**ity.- **institution** or the side of a river?- domestic **cats**, wild **cats**, or both.- **id** or errors.- genuine data **points** or errors.The tokens in MAX_ACTIVATING_TOKENS often represent specific items, categories, labels, or states that are being defined, differentiated, or described.The following tokens "or" and "," suggest these items are often presented as options, characteristics, or parts of a classification or explanation.Consider the examples:- "department or acting temporarily"- "branches, restoring a file, or creating a new branch"- "сельской местности, в городе или рядом с природой" (rural area, in the city or near nature)- "singular or plural"- "equal or not"- "institution or the side of a river"- "domestic cats, wild cats, or both"- "data points or errors"The patterns seem to be:1. **Keywords that define or categorize something**: department, branches, местности, singular, equal, institution, cats, id, points.2. These keywords often appear in contexts where **alternatives or specific attributes are being listed or contrasted**, frequently using "or" or commas to separate them.Let's try to synthesize this into a concise phrase (3-20 words).- "specific entities or states contrasted with alternatives" (7 words) - This seems good.- "categories or specific items presented as options" (7 words) - Also good.- "classification keywords followed by choices" (5 words) - A bit more direct.- "defining distinct items or states from others" (7 words) - Captures the essence.Let's review the prompt requirements:- Concise (3-20 words).- Captures what the neuron detects/predicts by finding patterns.- Find patterns in lists.- Single short phrase, not a full sentence.- Avoid "words related to", "concepts related to", "variations of the word".- Do not start with "This neuron detects/predicts".- Do not mention "tokens" or "patterns".- Do not capitalize first letter unless proper noun.- Specific.- Majority should match pattern.The phrase "specific entities or states contrasted with alternatives" fits well."categories or specific items presented as options" also fits."classification keywords followed by choices" is very direct and short.Let's re-evaluate the MAX_ACTIVATING_TOKENS.- department: a division or section- branches: divisions of a company; paths in a tree/graph- местности: geographical areas- singular: one item- equal: identical- in: preposition indicating inclusion or location- institution: an established organization- cats: a type of animal- id: identifier- points: marks, scores, or locationsThese are all nouns or adjectives that represent distinct concepts or identifiers.The TOKENS_AFTER_MAX_ACTIVATING_TOKEN are mostly "or" and ", ". These strongly suggest a choice, a list, or a separation between items.The TOP_ACTIVATING_TEXTS confirm this:- "department or is acting temporarily" -> contrast/choice- "switching branches, restoring a file, or creating a new branch" -> list of actions/states- "в сельской местности, в городе или рядом с природой" -> options for location- "singular or plural" -> contrast/choice- "equal or not" -> contrast/choice- "financial institution or the side of a river" -> contrast/choice- "domestic cats, wild cats, or both" -> list/choices- "Outliers: genuine data points or errors." -> contrast/choiceThe core idea is a *distinction* or a *choice* between specific items or states.Let's refine the phrase:- "distinguishing specific items or states" (5 words) - Does not capture the "or/," aspect well.- "items or states contrasted with alternatives" (6 words) - Good.- "specific classifications or conditions with alternatives" (6 words) - Also good.- "contrasting specific categories or states" (5 words) - Misses the "or" aspect.The phrase "specific items or states contrasted with alternatives" seems to best capture the tokens in MAX_ACTIVATING_TOKENS (items/states) alongside the follow-up tokens and the context from TOP_ACTIVATING_TEXTS (contrasted with alternatives).Let's check token counts if I were to try something derived from the TOP_POSITIVE_LOGITS.- every: e.g. "every department", "every branch"- dozens: "dozens of cats", "dozens of points"- COMDAT, also, etc are less clear.The pattern seems stronger in relation to MAX_ACTIVATING_TOKENS and TOKENS_AFTER_MAX_ACTIVATING_TOKEN combined with the activating texts.The phrase "specific items or states contrasted with alternatives" is 6 words, within the 3-20 word limit.It avoids forbidden phrases.It is specific.It describes what the neuron might be detecting based on the provided data.Let's consider: "defining items with contrasting options". (

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_32_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     veya
    0.64
     किंवा
    0.60
    或
    0.59
     বা
    0.58
    หรือ
    0.57
     или
    0.56
     hoặc
    0.55
     vagy
    0.54
    或者
    0.53
     அல்லது
    0.52
    POSITIVE LOGITS
    every
    0.39
     ايضا
    0.39
    COMDAT
    0.38
     dozens
    0.38
    ~.,
    0.38
     এরপর
    0.37
    Bigg
    0.37
     তারাও
    0.36
    ETON
    0.36
    もあった
    0.36
    Activations Density 0.007%

    No Known Activations