© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 11-GEMMASCOPE-2-TRANSCODER-262K
    4. 92729
    Prev
    Next
    INDEX
    Explanations

    The list of `MAX_ACTIVATING_TOKENS` clearly points to names and terms from Greek mythology: "Hym", "phone" (likely part of Aphrodite/telephone), "rodite" (Aphrodite), "omer" (Homer), and "goddess".The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` shows 'ns' after 'Hym', forming 'Hymns'.The `TOP_ACTIVATING_TEXTS` confirm this, mentioning "Homeric Hymns", "gods and goddesses", "Demeter", "Persephone", "Aphrodite", "Homer", "Iliad", "Odyssey", "Athena", "Poseidon".The `TOP_POSITIVE_LOGITS` list includes "любовь" (love in Russian), which is a core aspect of Aphrodite. It also includes Greek 'онів' (ions) and words that sound Greek like 'мия'.Combining these, the neuron is clearly interested in Greek mythology, particularly figures like Homer, Aphrodite, and goddesses, and concepts like hymns.A concise explanation could be: **Greek goddesses and Homeric hymns**Let's check the word count: 5 words. This is within the 3-20 word limit.It is specific.It avoids forbidden phrases.It doesn't capitalize 'greek' as it's not a proper noun in this context as the start of the phrase, and neither are 'goddesses' or 'homeric'.Greek goddesses and Homeric hymns

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_11_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    new
    0.58
    uz
    0.57
    Ver
    0.56
    Cum
    0.56
    IF
    0.55
     bete
    0.55
    Ze
    0.55
     st
    0.55
     
    0.55
    estado
    0.54
    POSITIVE LOGITS
     deployed
    0.68
     deploying
    0.66
     deployments
    0.61
     любовь
    0.60
    өн
    0.59
    ність
    0.58
    ्स
    0.57
     curtailed
    0.57
    онів
    0.57
    мия
    0.56
    Activations Density 0.006%

    No Known Activations