© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B
    3. 31-GEMMASCOPE-2-RES-262K
    4. 168799
    Prev
    Next
    INDEX
    Explanations

    medical or research limitations

    np_acts-logits-general · gemini-2.5-flash-lite

    thoughtsThe user wants me to explain the behavior of a neuron.The explanation must be a concise phrase (3-20 words).I need to identify a pattern from the four provided lists: `TOP_POSITIVE_LOGITS`, `TOP_ACTIVATING_TEXTS`, `MAX_ACTIVATING_TOKENS`, and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`.Let's analyze the lists:`MAX_ACTIVATING_TOKENS`:`how`, `classics`, `currently`, `Off`, `orough` (likely thorough), `current`, `'`, `feel``TOKENS_AFTER_MAX_ACTIVATING_TOKEN`:`giant`, `model`, `*`, `remain`, `have`, `"`, `*`, `bi``TOP_POSITIVE_LOGITS`:`buying`, `viral`, `eczema`, `धपुर` (special char, ignore), `alleviate`, `liabilities`, `hurd` (likely hurt), `viral`, `ETING` (likely marketing), `netted``TOP_ACTIVATING_TEXTS`:- "But...giant robots!"- "understand my challenges, I have prepared some information Patient : What if they reject me?"- "explains *how* gravity works"- "Unique & Off-the-Beaten-Path: While classics remain popular"- "where you currently have the `data/coco` directory."- "modes: "Heater Off" and "Heater On"."- "Be *thorough* here."- "Beyond the current biomes"- "* Let's unpack these feelings."- "my resume and cover letter as honest as they get."Observations:1. **`MAX_ACTIVATING_TOKENS`**: Contains words like `how`, `classics`, `currently`, `Off`, `thorough`, `current`, `feel`. These seem like descriptive words, states, or question starters.2. **`TOKENS_AFTER_MAX_ACTIVATING_TOKEN`**: Contains words like `giant`, `model`, `remain`, `have`, `bi`.3. **`TOP_POSITIVE_LOGITS`**: Contains words like `buying`, `viral`, `alleviate`, `liabilities`, `hurd`/`hurt`, `netted`. These seem related to outcomes, problems, or summaries.4. **`TOP_ACTIVATING_TEXTS`**: * Shows examples like "giant robots!", "how gravity works", "classics remain popular", "currently have", "Heater Off", "thorough here", "current biomes", "feelings". * The texts often describe present states, conditions, or explanations. * Some texts involve descriptions like "giant robots", "classics remain popular", "current biomes", "Heater Off/On".Let's try to find a common theme or pattern that connects these lists.Focusing on `MAX_ACTIVATING_TOKENS` and `TOP_ACTIVATING_TEXTS`:- `how` -> "how gravity works"- `classics` -> "classics remain popular"- `currently` -> "currently have"- `Off` -> "Heater Off"- `orough` (thorough) -> "Be *thorough* here."- `current` -> "current biomes"- `feel` -> "feelings"These seem to be about describing states, conditions, or providing explanations/definitions. The `TOP_POSITIVE_LOGITS` (`buying`, `viral`, `alleviate`, `liabilities`, `hurt`, `netted`) might represent outcomes or evaluations *related to* these states or descriptions. For example, an explanation of how something works might lead to a decision about buying something, or how to alleviate a problem.The pattern seems to be about describing the *current state*, *how things are*, *specific conditions*, or *features*.Let's consider the structure:`MAX_ACTIVATING_TOKENS` often come from the middle or end of sentences, describing a state or attribute.`TOKENS_AFTER_MAX_ACTIVATING_TOKEN` seem to be continuations, but not always directly related in a simple syntactic way.A key aspect that seems to emerge is the nomination of *specific states* or *conditions*.- "how gravity works" (explaining a mechanism)- "classics remain popular" (describing a lasting trend/state)- "currently have" (describing a present state)- "Heater Off" (describing a state)- "Be *thorough*" (describing a quality of action)- "current biomes" (describing a present characteristic)- "feelings" (describing an internal state)The `TOP_POSITIVE_LOGITS` provide clues to what kind of concepts are associated with these states. `viral`, `alleviate`, `liabilities`, `netted` suggest contexts where conditions or states might lead to actions or outcomes.Let's look at `MAX_ACTIVATING_TOKENS` again:`how`, `classics`, `currently`, `Off`, `orough`, `current`, `'`, `feel`And `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`:`giant`, `model`, `*`, `remain`, `have`, `"`, `*`, `bi`If we look at the sample texts:- "*continues watching anime*" (context of distraction/feeling)- "understand my challenges, I have prepared some information" (context

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-pt/resid_post/layer_31_width_262k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     simbol
    0.53
    Produto
    0.48
    ่
    0.46
    3
    0.46
     символ
    0.46
     Facultad
    0.45
    Lantern
    0.43
    4
    0.43
    Pacific
    0.43
     Symbol
    0.43
    POSITIVE LOGITS
     buying
    0.51
    弛
    0.49
     quarantined
    0.46
    ETING
    0.45
     distressing
    0.45
     childcare
    0.45
    hass
    0.45
     alleviate
    0.44
     hurts
    0.44
     crippling
    0.44
    Activations Density 0.000%

    No Known Activations