© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-12B
    3. 24-GEMMASCOPE-2-RES-16K
    4. 926
    Prev
    Next
    INDEX
    Explanations

    The provided `MAX_ACTIVATING_TOKENS` list is empty.Therefore, I cannot identify the maximum activating token to use as a fallback according to the instructions.However, let's look at `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`:- `very`- `all`And `TOP_ACTIVATING_TEXTS`:- "It's a *very* active field"- "Winning typically requires matching *all* the numbers drawn."And `TOP_POSITIVE_LOGITS`:- `ad`- `ot`- `acceler`- `ancien`- `évaluation`The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` list contains "very" and "all". These words appear in the `TOP_ACTIVATING_TEXTS` associated with this neuron.These words suggest endings or completions. "very" often modifies adjectives (e.g., "very active"). "all" suggests totality or completeness (e.g., "all the numbers").The `TOP_POSITIVE_LOGITS` include words that might be related to specific domains or contexts.- `acceler` (likely related to 'acceleration' or 'accelerate')- `ancien` (French for 'old' or 'ancient')- `évaluation` (French for 'evaluation')- `ad` (adjective, advertisement)Given `very` and `all` as the primary indicators and their typical usage, they point towards an endpoint or a complete state.Considering the limitation of the empty `MAX_ACTIVATING_TOKENS` and the pattern in `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`, I will focus on the common theme implied by "very" and "all" when they appear at the end of segments or descriptions.The phrase needs to be 3-20 words, a phrase, specific, lowercase first letter, and avoid "tokens" or "patterns".Possible interpretations:- Focus on completion: "completion of phrases" or "end of descriptions"- Focus on intensity/completeness: "very and all" (too literal, not specific enough)Let's re-examine the `TOP_POSITIVE_LOGITS` for any unifying theme that might relate to "very" or "all".`ad`, `ot`, `acceler`, `ancien`, `évaluation`.These seem somewhat disparate.The most direct pattern comes from `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`: "very" and "all".These are often used to intensify or qualify something that has just been mentioned.- "very active field"- "matching all the numbers"The neuron might be looking for words that complete a thought or provide a definitive qualification.Let's try to capture the essence of "very" and "all" as qualifiers or completors.Option 1: Focus on "all" as totality."completing a set" (4 words)"all of something" (3 words)Option 2: Focus on "very" as intensity."intensifying a description" (3 words)Option 3: Combine them conceptually?"final qualifiers" (2 words - too short)"definitive qualifiers" (2 words - too short)Let's evaluate "completion of phrases". This implies the neuron is active when something is finished or finalized. "very" often comes after an adjective that might be the "end" of a descriptive phrase until "very" adds more info.completing a description

    np_acts-logits-general · gemini-2.5-flash-lite

    This neuron detects the underscores that mark the boundaries of italicized text (i.e. italic formatting).

    oai_token-act-pair · o4-miniTriggered by @jyhe0408
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-12b-pt/resid_post/layer_24_width_16k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     distaste
    0.65
     qualifies
    0.64
    icited
    0.64
     clearContext
    0.64
     arkadaş
    0.63
    ̷
    0.60
     …….
    0.59
    នុស្ស
    0.58
     Rashford
    0.57
     endIndex
    0.56
    POSITIVE LOGITS
    modus
    0.66
    Journal
    0.63
    L
    0.63
     Daily
    0.62
    Bloomberg
    0.60
    
    0.60
    bach
    0.59
    //
    0.59
    Oui
    0.59
    P
    0.57
    Activations Density 0.007%

    No Known Activations