© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 27-GEMMASCOPE-2-TRANSCODER-262K
    4. 117775
    Prev
    Next
    INDEX
    Explanations

    This neuron seems to be activated by words describing temperature sensations. Looking at the provided lists:* **MAX_ACTIVATING_TOKENS**: contains words like `hot`, `cold`, `cool`, `warm`.* **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: predominantly punctuation or connectors like `,`, `.`, `and`.* **TOP_POSITIVE_LOGITS**: includes words like `justamente` (precisely in Spanish), `autres` (others in French), `además` (besides/also in Spanish), suggesting cross-lingual or general descriptive contexts.* **TOP_ACTIVATING_TEXTS**: includes phrases like "It was cold, even in July", "Raves are *hot*", "The assembly hall was so hot and moist!", "runs incredibly hot, even with cooling sheets", "climate control went haywire, turning the open-plan office into a rotating cycle of arctic tundra and Saharan desert", "Temperature is usually slightly *too* cool", "utterly unsuitable for the sweltering August afternoon.", "The lecture hall was stifling.", "It is hot outside.temperature descriptors

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_27_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     संतुलित
    0.37
     inept
    0.36
     quid
    0.36
    阴
    0.36
     أرض
    0.35
     графи
    0.35
    рина
    0.34
    良好
    0.34
     स्मिथ
    0.34
    স্য
    0.34
    POSITIVE LOGITS
     justamente
    0.40
    টে
    0.39
     مثل
    0.38
     middlewares
    0.37
     ->
    0.36
    odoxy
    0.36
    zné
    0.36
    ጊ
    0.36
    autres
    0.36
     además
    0.36
    Activations Density 0.000%

    No Known Activations