© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 37-GEMMASCOPE-2-TRANSCODER-262K
    4. 205026
    Prev
    Next
    INDEX
    Explanations

    Plugin Name- `())` (after `sum`)Observations:1. `Content` followed by `Warning` or `Warnings` strongly suggests "Content Warning".2. `Plugin` followed by `Name` suggests "Plugin Name".3. `sum` appears twice. After the first `sum`, the token is `())`. After the second `sum`, the token is `())`. This might relate to `df.isnull().sum()` and `print(df.isnull().sum())` in the text, where `sum()` is used.4. `ality` is in `MAX_ACTIVATING_TOKENS`. The token after it in `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` is `Summary`. This is a bit of an outlier if we look at direct pairings. However, the *concept* might be that the neuron activates around things that are summaries or descriptions, possibly related to functionality or content. For example, "code functionality...".Let's look at the `TOP_ACTIVATING_TEXTS` for context.- " code functionality and give a threat score(1-10 low is safe)" -> mentions "functionality" and a score.- "## Code Functionality Summary" -> mentions "Functionality Summary".- "I've included a content warning" -> "content warning".- "**Content Warning:** This story contains themes..." -> "Content Warning".- "Important!) # Check for missing values and handle them. This is crucial. print(df.isnull().sum())" -> `sum()` is used.- "Please list alternatives to kubernetes" -> mentions "alternatives".- "**Content Warning:** This poem deals with themes..." -> "Content Warning".- "I will also include content warnings at the beginning." -> "content warnings".- "/** * Plugin Name: My Awesome Plugin */" -> "Plugin Name".- "/** * Plugin Name: MyPlugin */" -> "Plugin Name".- "print(df.isnull().sum())" -> `sum()` is used.Strong patterns emerging:1. **Content Warning**: `Content` + `Warning`/`Warnings`2. **Plugin Name**: `Plugin` + `Name`3. **Summation/Summary**: `sum` (related to code/data checks) and potentially `ality`/`Summary` (related to summaries of content/functionality).4. **Alternatives**: Mentioned in text and `Alternatives` token.The neuron seems to be sensitive to specific keywords that indicate descriptive or functional elements

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_37_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    详解
    0.52
     analysis
    0.46
     explained
    0.46
     comparison
    0.45
    分析
    0.45
     appreciation
    0.44
     ode
    0.44
    一覧
    0.43
     revealed
    0.42
    介紹
    0.41
    POSITIVE LOGITS
    㞰
    0.41
     Сасик
    0.38
    岧
    0.38
    Hide
    0.37
    ぅ
    0.37
    masını
    0.36
    Casey
    0.36
    掳
    0.36
    㞖
    0.36
     মিনিটের
    0.35
    Activations Density 0.003%

    No Known Activations