© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 27-GEMMASCOPE-2-TRANSCODER-262K
    4. 120023
    Prev
    Next
    INDEX
    Explanations

    The neuron seems to detect patterns where specific words or phrases indicate a continuation, definition, or specific action/state.Let's break down the provided lists:1. **MAX_ACTIVATING_TOKENS:** This list is truncated in the prompt ("..."), making it hard to directly analyze the activating tokens themselves. However, its presence alongside the other lists is crucial.2. **TOKENS_AFTER_MAX_ACTIVATING_TOKEN:** `fas`, `the`, `states`, `can`, `to`, `a`, `CONTRACT`, `nearly`, `cont`, `use` These are tokens that often *follow* whatever the MAX_ACTIVATING_TOKENS are. Looking at these, some seem like prefixes or parts of words: * `fas` -> suggests `fascinating` (seen in TOP_ACTIVATING_TEXTS) * `states` -> `states` (seen in TOP_ACTIVATING_TEXTS) * `can` -> `can be found` (seen in TOP_ACTIVATING_TEXTS) * `CONTRACT` -> `CONTRACTOR` (seen in TOP_ACTIVATING_TEXTS) * `nearly` -> `nearly half` (seen in TOP_ACTIVATING_TEXTS) * `cont` -> `contributes` (seen in TOP_ACTIVATING_TEXTS) * `use` -> `use evidence` (seen in TOP_ACTIVATING_TEXTS)3. **TOP_POSITIVE_LOGITS:** `urit`, `म्मद`, `NotBlank`, `]`, `gravid`, `apeutic`, `Cock`, `breast`, `[`, `]`, `uracy` These are the words/tokens the neuron is *most associated with*. * `NotBlank`: A programming term, implies a check or condition. * `gravid`: Medical term related to pregnancy. * `breast`: Anatomical term, often related to health. * `therapeutic`: Medical/treatment related. * `uracy`: Likely part of a word like `accuracy` or `curacy`. These suggest the neuron might be looking for specific concepts, possibly domain-specific (medical, technical).4. **TOP_ACTIVATING_TEXTS:** * "...fascinating, involving character study" * "...states in the North have utilised their diplomatic influence..." * "...can be found in some literature." * "...CONTRACTOR had duly submitted the documents..." * "...nearly half of all social media users arc likely to unfollow..." * “…contributes to a more nuanced understanding of ECG filtering…” * "...use evidence and childhood curiosity..."**Pattern Identification:*** **Continuation/Definition:** The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` show common continuations. Phrases like "...can be found...", "...states...", "...use evidence...", "...contributes...", and "...nearly half..." suggest the neuron might be identifying phrases that link concepts, provide definitions, or state facts/conditions.* **Specific Terms:** `TOP_POSITIVE_LOGITS` points to specific, sometimes technical or domain-specific terms (e.g., `NotBlank`, `gravid`, `therapeutic`, or related to `accuracy`).* **Contextual Links:** The activating texts show how these tokens are used (e.g., linking character study to fascination, diplomatic influence to states, common findings to literature, contractor actions, user behavior to unfollowing, contributions to understanding, and evidence to curiosity).The neuron appears to be capturing the *contextual linking in phrases*, especially around stating facts, conditions, or contributions, potentially with a bias towards specific domain terms.Considering the prompt: "find a pattern that helps you explain the behavior of the neuron." and "concise explanation (3 to 20 words) that captures what the neuron detects or predicts by finding patterns in lists." and "say the pattern itself".The most prominent pattern is how certain tokens signal a specific kind of statement or elaboration. Phrases starting with "can be found", "states", "nearly", "contributes", "use" all fall under explaining or describing. The `TOP_POSITIVE_LOGITS` suggest some specific themes (medical, technical checks). However, the *pattern* is more about the *structure* of these informative sentences."Identifying informative phrase continuations" is too long."Explaining conditions or facts" captures some of it."Linking concepts or stating facts" is also good.Let's look at the `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` again: `fas`, `the`, `states`, `can`, `to`, `a`, `CONTRACT`, `nearly`, `cont`, `use`.These are often part of descriptive or declarative sentences.The `TOP_POSITIVE_LOGITS` like `NotBlank`, `gravid`, `therapeutic`, `accuracy` point to specific *types* of information.Combining these, the neuron might be looking for sentences that *describe, qualify, or specify*.Consider "states" and "can be found" and "nearly half" and "contributes" and "use evidence". These are all descriptive.Let's try to synthesize into a short phrase:- "States facts or provides clarification" (6 words)- "Describes conditions or actions taken" (5 words)- "Links concepts to states or actions" (6 words)- "Clarifying phrases and descriptive statements" (5 words)Given the presence of `NotBlank`, `gravid`, `therapeutic`, `accuracy` (from `TOP_POSITIVE_LOGITS`), the neuron might be leaning towards technical or medical specifications/clarifications. But the `TOP_ACTIVATING_TEXTS` are broader than just medical.The phrase "states facts or provides clarification" seems to cover the variety shown."states" token: "states in the North have utilised...""can": "...can be found in some literature." (clarification)"nearly": "...nearly half of all social media users..." (fact/statistic)"cont": "...contributes to a more nuanced understanding..." (clarification/explanation)"use": "...use evidence and childhood curiosity..." (action/method)The prompt also mentions "finding patterns in lists". The pattern is in *how* language is used to convey information.Let's re-evaluate the core task: "concise explanation (3 to 20 words) that captures what the neuron detects or predicts by finding patterns in lists."The neuron is detecting language that elaborates, specifies, or states things.What about phrases related to *definitions or specifications*?- "Defining terms and specifying conditions" (5 words)- "Specifying conditions or stating facts" (5 words)Looking at `TOP_POSITIVE_LOGITS` again: `NotBlank` (condition/spec), `gravid` (medical spec), `therapeutic` (medical spec), `accuracy` (spec). This strongly hints at specification or condition-setting.The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` (`states`, `can`, `nearly`, `cont`, `use`) are often *part of* such specifications or facts."Specifying conditions or stating facts" seems quite fitting.Let's consider an alternative angle: the neuron is activated when a *statement is made that needs further qualification

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_27_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     படங்கள்
    0.29
    ხვ
    0.28
     pitan
    0.27
     प्रश्न
    0.27
     outros
    0.27
     lainnya
    0.27
    лом
    0.26
    )**
    0.26
    elizmente
    0.26
     речь
    0.26
    POSITIVE LOGITS
    urit
    0.33
    म्मद
    0.33
    NotBlank
    0.32
    `],
    0.32
     gravid
    0.32
    apeutic
    0.32
    Cock
    0.31
    breast
    0.31
    "+"|".
    0.31
    uracy
    0.30
    Activations Density 0.001%

    No Known Activations