© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 15-GEMMASCOPE-2-TRANSCODER-262K
    4. 14241
    Prev
    Next
    INDEX
    Explanations

    **I am analyzing the provided lists to identify patterns and predictive functions of the neuron.****MAX_ACTIVATING_TOKENS**: `for` (repeatedly)**TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: `addressing`, `the`, `Reactions`, `calls`, `a`, `this` (among others)**TOP_POSITIVE_LOGITS**: `ino`, `ina`, `ini`, `3`, `inu`, `q`, `ياه`, `iana`, `idet`, `inos` (These seem to be subword tokens or possibly foreign language parts, indicating the neuron might be sensitive to certain phonetic structures or beginnings/endings of words.)**TOP_ACTIVATING_TEXTS**:- "AWS Activate credits. It's not a general next step. Therefore, the best course of action is to prepare for **addressing** AWS-related issues."- "Offer Contingency Plans: We'll prepare for the unexpected."- "Prepare for **Reactions**: Anticipate potential **reactions** (anger, sadness, denial, shock) and think about how you'll"- "preparing **for calls**."- "preparing for the future"- "prepare for a Counteroffer"- "preparing for this."- "preparing for this, studying Martian soil samples"- "Preparing for 'empty' islands is crucial for a pirate-hunt campaign."**Analysis**:The `MAX_ACTIVATING_TOKENS` are consistently "for

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_15_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     fitur
    0.52
    㹮
    0.50
     conceptually
    0.49
     fonctionnalités
    0.49
    𝔞
    0.49
    PrototypeOf
    0.48
     geranium
    0.47
    гка
    0.47
     grafts
    0.46
     grits
    0.46
    POSITIVE LOGITS
    ino
    0.59
    ina
    0.57
    ini
    0.54
    3
    0.52
    inu
    0.51
    q
    0.50
    ياه
    0.48
    iana
    0.48
    idet
    0.48
    inos
    0.47
    Activations Density 0.011%

    No Known Activations