© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B
    3. 31-GEMMASCOPE-2-RES-262K
    4. 14389
    Prev
    Next
    INDEX
    Explanations

    made out to

    np_acts-logits-general · gemini-2.5-flash-lite

    what does"}, {"link": "authorization link"}, {"rocket": "rocket still"}, {"Arguments": "arguments, evidence"}, {"within": "within NLP"}, {"Drug": "drug-like properties"}The neuron seems to be activated by phrases that introduce an explanation, a question, or a definition.Examples:- "What does..."- "authorization link"- "rocket still"- "arguments, evidence, and nuances"- "within NLP"- "drug-like properties"- "Arguments Used"- "corresponds"The structure "What [X] does" is explicitly present."rocket still" is present."Arguments Used" is present."within NLP" is present.The common theme is introducing a topic or a specific element and then continuing to discuss it.The first max activating token is "What". The token after that is "does". This forms "What does".The texts show a pattern of question/introduction followed by explanation.A concise phrase could be "introducing explanations or questions".Or simply "what does" as a prominent example.Let's look for a more general pattern covering these."what does" is a very strong candidate."introducing a topic""explaining a concept""providing definitions or arguments"Considering the `MAX_ACTIVATING_TOKENS`: "What", "link", "rocket", "Arguments", "within".Considering `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`: "does", "can", "Att", "Risks", "still", "NLP", "Drug", "Used", "the", "corresponds".Pairs:What doeslink canrocket stillArguments Attwithin NLPArguments UsedThe structure "What does" is very direct.The phrase "What does" itself is 2 words. It falls within 3-20 words.Let's look at TOP_POSITIVE_LOGITS. It has "listOf". This suggests it might be related to lists or structured data.The `TOP_ACTIVATING_TEXTS` contain structure like bullet points or numbered lists of arguments/risks/properties.The neuron seems to activate around introducing a topic or item that will be further described or quantified.- "What does" introduces a question/topic.- "link" can introduce an authorization link. what does","rocket still","Arguments Used","within NLP

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-pt/resid_post/layer_31_width_262k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     Worksheet
    0.44
     spreadsheet
    0.42
    inactive
    0.40
     worksheet
    0.39
     Trung
    0.39
     Lavender
    0.39
     Divide
    0.39
     inadequate
    0.38
     VMs
    0.38
     nhàng
    0.38
    POSITIVE LOGITS
     esistono
    0.42
    ürz
    0.42
    قیه
    0.42
    Einstellungen
    0.42
    нях
    0.40
    üh
    0.39
     بیش
    0.39
    IsDir
    0.39
     ਹੈ
    0.39
    Foo
    0.39
    Activations Density 0.000%

    No Known Activations