© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-4-31B
    3. 30-RES-MATRYOSHKA-131K
    4. 128448
    Prev
    Next
    INDEX
    Explanations

    This neuron is analyzing spatial relationships, particularly how individuals or objects are positioned relative to each other or a reference point. Common themes include "sitting," "around," "on," "back," and "flanked," suggesting a focus on arrangement, proximity, and orientation. The subsequent tokens like "the," "in," "and," "different," "top," and "by" often follow these positional indicators, helping to describe the scene or context further.Looking for a concise phrase (3-20 words) that captures this.Keywords found:- sitting- around- circle- on- back- flanked- to- at- approximatelyThese all relate to positions, locations, and arrangements.Let's look at TOP_POSITIVE_LOGITS for reinforcing context:- makeshift- física- on- vreau-小编- egun- Acrylic- solitude- impromptu- Safety"física" (Spanish for physical) and "Acrylic" (material) could relate to descriptions of objects or scenes. "makehsift", "impromptu" suggest arrangement. "solitude" might imply being positioned alone. "on" is repeated.Let's look at TOP_ACTIVATING_TEXTS to see the context of MAX_ACTIVATING_TOKENS and TOKENS_AFTER_MAX_ACTIVATING_TOKEN.This neuron detects spatial arrangements and positioning

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    decoderesearch/gemma-4-saes/gemma-4-31b
    Prompts (Dashboard)
    16,384 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     again
    -0.07
     다시
    -0.07
     for
    -0.07
    ට්ට
    -0.07
    再次
    -0.07
     ponieważ
    -0.06
     ponownie
    -0.06
     nuovamente
    -0.06
     wieder
    -0.06
     returned
    -0.06
    POSITIVE LOGITS
     makeshift
    0.08
     física
    0.06
     on
    0.06
     vreau
    0.06
    小编
    0.06
     egun
    0.06
    Acrylic
    0.06
     solitude
    0.06
     impromptu
    0.06
    Safety
    0.06
    Activations Density 0.005%

    No Known Activations