© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-12B
    3. 24-GEMMASCOPE-2-RES-16K
    4. 898
    Prev
    Next
    INDEX
    Explanations

    Okay, let's break down the pattern observed in the provided lists to explain the neuron's behavior.- **MAX_ACTIVATING_TOKENS**: We see punctuation marks like `;` and `?`.- **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: We see `<end_of_turn>`, `<start_of_turn>`, `d`, and `I`. This suggests the neuron is looking at what happens *after* certain punctuation or turn boundaries, potentially recognizing continuations or distinct speakers/sections.- **TOP_POSITIVE_LOGITS**: This includes words like `and`, `by`, and also non-English words (`مرين`, `اور`, `ش`, `তাহাই`). This heterogeneity is a bit tricky.- **TOP_ACTIVATING_TEXTS**: The texts include SQL commands (`grant select on all`), explanations of terms (`mean log pass rate`), questions, and responses. There are also English phrases and some non-English characters within or near the output. A key observation is the transition between user input and model output, often marked by `<start_of_turn>` and `<end_of_turn>` markers. The presence of `?` and `;` in MAX_ACTIVATING_TOKENS and then `<end_of_turn>` or sequences like `d<start_of_turn>I<start_of_turn>` afterwards suggests the neuron might be distinguishing between different types of turns or statements, perhaps recognizing the end of a user's query or a signal for the model to start.The phrase "transitioning between turns" seems too long or descriptive.The presence of punctuation like `;` and `?` followed by turn markers suggests the neuron might be identifying boundaries or specific types of content within conversational turns.Looking at the `TOP_POSITIVE_LOGITS`, they are quite varied.The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` show `<end_of_turn>`, `<start_of_turn>`, and single letters.Let's reconsider the pattern.The neuron seems to activate when encountering punctuation (`?`, `;`) and then subsequently processing tokens that signify the end of one turn and potentially the beginning of another (`<end_of_turn>`, `<start_of_turn>`). The `TOP_ACTIVATING_TEXTS` show this is common in conversational data where turns are delimited.A plausible hypothesis is that this neuron detects the boundary between a user query/statement and a model response, or more generally, the end of a conversational segment. The punctuation might be a signal *for* this boundary.Let's try to find a concise phrase (3-20 words) that captures this."End of user query" - too specific to user."Turn boundary detection" - uses "pattern" implicitly."End of statement" - a bit general.Consider the combination of `?`, `;` and `end_of_turn`, `start_of_turn`.The texts show distinct turns.The neuron might be recognizing the transition point or the markers themselves.Let's look at the *meaning* derived from these.The SQL examples show a complete command ending in `;`.The "mean log pass rate" example ends with a question mark `?`.The multiple-choice example ends with `d. All of them<end_of_turn>`.What follows `;` or `?` are often turn beginnings/endings.This neuron seems to be related to parsing conversational turns or structured input endings.If we focus just on the direct findings:MAX_ACTIVATING_TOKENS: `;`, `?`TOKENS_AFTER_MAX_ACTIVATING_TOKEN: `<end_of_turn>`, `<start_of_turn>`, `d`, `I`This strongly suggests it's about the structure of conversation turns and how they end or begin.Let's try to find a common theme. The neuron seems to identify the *point* where a piece of information (like a query, a command, or a response fragment) concludes, often preparing for the next segment.How about: "detecting statement endings" or "boundary markers"?The prompt asks for *what the neuron detects or predicts by finding patterns in lists*.The pattern is punctuation (`?`, `;`) followed by turn delimiters or the start of a new turn.What about focusing on the *type* of ending?The SQL command is precise. The question is asking for information. The multiple choice answer is a conclusion.Let's look at the `TOP_POSITIVE_LOGITS` again: `and`, `alias`, `by`. These might hint at structural elements or relationships.If the neuron activates on `;` or `?` and then sees turn markers, it's essentially identifying structural segmentation in text.Consider the examples:- `grant select on all to sw_dc_da;` -> `;` followed by model response start.- `Which team have the highest chance to win between varnamo and varbergs?` -> `?` followed by model response start.- `d. All of them<end_of_turn>` -> This is a model response ending.The prompt is *not* to list what follows, but what the neuron *detects*.The neuron detects the end of declarative or interrogative statements followed by conversational turn changes.The actual tokens are structured text elements indicating conversational flow.Let's try to simplify.It's processing the structural elements of dialogue and structured text.statement and turn endings

    np_acts-logits-general · gemini-2.5-flash-lite

    The neuron fires on mentions of JavaScript filenames or paths—i.e. tokens naming “*.js” module files.

    oai_token-act-pair · o4-miniTriggered by @jyhe0408
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-12b-pt/resid_post/layer_24_width_16k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     degrés
    0.49
     Τ
    0.49
     τα
    0.47
     България
    0.46
     quelconque
    0.46
    याच्या
    0.46
     trouv
    0.46
     tieto
    0.46
     meningkatkan
    0.46
     EBIT
    0.46
    POSITIVE LOGITS
    clickHandler
    0.55
    algo
    0.54
    navigationItem
    0.51
    alias
    0.50
    js
    0.50
    webview
    0.49
    handlebars
    0.48
    сшта
    0.47
    
    0.47
    getInstance
    0.47
    Activations Density 0.001%

    No Known Activations