© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B
    3. 31-GEMMASCOPE-2-RES-262K
    4. 237593
    Prev
    Next
    INDEX
    Explanations

    explanation of the neuron's behavior.The neuron appears to activate for descriptions involving specific terminology, particularly those found in technical contexts, academic writings, or specific industry jargon. This is suggested by the presence of words like 'eclectic' and 'means' from `MAX_ACTIVATING_TOKENS`, followed by general terms in `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`, and a mix of specific, sometimes foreign, concepts in `TOP_POSITIVE_LOGITS` (like "तूफान", "Capsule", "magnificence", "glider", "hanger", "飞", "আকাশ"). The `TOP_ACTIVATING_TEXTS` further support this with phrases like "eclectic in his use of these sources", "means it's an addition to your diet", "base file for the particular model", and "you mentioned".This neuron seems to capture a transition from a general contextual word to a more specific descriptor or technical term. It might be about defining or specifying something by its properties or components.Given the instructions, I need to find a concise pattern.Looking at `MAX_ACTIVATING_TOKENS`: 'to', 'Your', 'eclectic', 'means', 'file', 'you'.Looking at `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`: 'Rooms', 'Located', 'A', 'There', 'blame', 'teenager', 'in', 'it', 'for', 'mentioned'.Let's re-evaluate the `TOP_POSITIVE_LOGITS`. Many of these are foreign words or specific items:- तूफान (storm/typhoon)- Capsule- magnificence- ף (likely noise, but could be special character)- আকাশ (sky)- glider- திட (sudden/solid)- 飞 (fly)- hanger- বহ (many/much)These suggest topics like weather, transportation, specific objects, or abstract qualities.Now look at `TOP_ACTIVATING_TEXTS` again.- "Rooms come with a flat-screen TV with satellite channels..." (describing features of rooms)- "Located within a 9-minute walk of Sochi Winter Theatre..." (describing location)- "A 2 star property approximately 4.66 km (2.91 mi) from Cala Llonga" (describing property rating and distance)- "There are 2 private institutes near BOULEVARD 88." (describing proximity)- "base file for the particular model." (describing technical configuration)- "you mentioned that are not responding to your WiFi LAN." (describing technical issue)- "This is a 'supplement', means it's an addition to your diet..." (defining a term)- "He is eclectic in his use of these sources." (describing methodology)- "blame earthquakes on 'sinfulness...'"- "Your teenager is evasive..."The `MAX_ACTIVATING_TOKENS` like "file", "means", "Your" and the subsequent tokens like "blame", "teenager", "mentioned" alongside the diverse `TOP_POSITIVE_LOGITS` suggest a neuron that might be identifying technical terms, definitions, or specific properties/attributes being discussed.It seems to be about listing or describing specific items, features, or concepts.Let's consider the pattern "word + definition/description".- "means it's an addition..."- "file for the particular model."- "Your teenager is evasive..."The `TOP_POSITIVE_LOGITS` being very diverse suggests the neuron isn't tied to a single concept but rather a *type* of information.Could it be about specifying things?"Your teenager", "base file", "A 2 star property".The combination of "means" in `MAX_ACTIVATING_TOKENS` and the definition explanation in `TOP_ACTIVATING_TEXTS` is strong.The "file" token is followed by "for the particular model", also a descriptive phrase."Your" followed by "teenager" or "eclectic in his use" describes a person's trait or state.Let's look at the `TOP_POSITIVE_LOGITS` again. They are very eclectic, spanning languages and concepts. This points away from detecting specific topics and towards a structural pattern.The pattern seems to be: a common word, followed by a specific descriptor or definition.Consider the structure `MAX_ACTIVATING_TOKEN` + `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`.- 'to' + 'Rooms' (describing room features)- 'Your' + 'Located' (describing a place)- 'eclectic' + 'A' (describing something)- 'means' + 'There' (introduces existence/description)- 'file' + 'blame' (less clear here)- 'you' + 'teenager' (describing a person)- 'you' + 'mentioned' (referring to something specific)The `TOP_POSITIVE_LOGITS` are very diverse: specific objects (Capsule, glider, hanger), abstract concepts (magnificence), foreign words that often represent specific things/concepts (तूफान, आकाश, திட, 飞), and grammatical particles/affixes (ף, বহ).This points to the neuron activating on descriptions, definitions, or specific components/attributes.Let's try to formulate a phrase."Defining specific attributes or components" - too long."Describing specific information" - a bit vague."Explaining definitions or components" - also too long.The neuron seems to activate when information is being specified or defined.Words like 'means', 'file', 'your' in MAX_ACTIVATING_TOKENS often lead to specificity in the text.The TOKENS

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-pt/resid_post/layer_31_width_262k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     composés
    0.43
    s
    0.43
    issaient
    0.42
     словарь
    0.41
    についても
    0.41
    البات
    0.41
    valuation
    0.39
    dit
    0.39
    riture
    0.38
    ssä
    0.38
    POSITIVE LOGITS
     तूफान
    0.46
     Capsule
    0.46
     magnificence
    0.45
    ף
    0.45
     আকাশ
    0.45
     glider
    0.45
     திட
    0.44
    飞
    0.43
     hanger
    0.43
    বহ
    0.42
    Activations Density 0.000%

    No Known Activations