© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 31-GEMMASCOPE-2-RES-262K
    4. 81824
    Prev
    Next
    INDEX
    Explanations

    most artistic", "more extensive", "more philosophical", "most artistic", "more maneuverable" suggests a focus on comparative or superlative adjectives describing a quality or degree.The presence of "most", "more", and "eless" in MAX_ACTIVATING_TOKENS also points to this comparative/superlative theme.Looking at "TOKENS_AFTER_MAX_ACTIVATING_TOKEN":- 'artistic' after 'most'- 'extensive' after 'more'- 'artistic' after 'most'This strongly suggests a pattern related to comparative and superlative forms of adjectives.Proposed Explanation: comparative or superlative adjectives.Let's check length: 4 words. This fits the criteria.Is it specific? Yes.Does it avoid forbidden phrases? Yes.Does it start with a lowercase letter? Yes.Consider other tokens:'Values' -> followed by ':''years' -> followed by '.''license' -> followed by '*''mindset' -> followed by '),'These don't fit the comparative/superlative pattern as well, but the 'most'/'more'/'artistic'/'extensive' pattern is very strong in both MAX_ACTIVATING_TOKENS and TOP_ACTIVATING_TEXTS.Let's refine: "comparative or superlative adjectives" might be too broad.The examples are: 'most artistic', 'more extensive', 'more philosophical', 'more maneuverable'. These are intensifiers followed by adjectives.How about " Intensifier + adjective "?'most artistic', 'more extensive', 'more philosophical', 'more maneuverable'.This seems quite accurate. 'most', 'more' are intensifiers.Let's test the word count. "intensifier adjective" is 2 words. Too short."intensifier followed by adjective" is 4 words."intensifiers like most and more" is 5 words."describing degrees or comparisons" is 4 words. This captures the essence well.Let's re-evaluate the tokens:MAX_ACTIVATING_TOKENS: most, more, eless, Values, years, license, most, ,, mindsetTOKENS_AFTER_MAX_ACTIVATING_TOKEN: artistic, extensive, ness, :**, ., *, artistic, *, relying, ),TOP_ACTIVATING_TEXTS:- "most artistic"- "more extensive"- "more philosophical"- "Values: "- "seven years."- "food license"- "most artistic", "more maneuverable"- "mindset"The most consistent pattern is definitely around intensification and comparison."most artistic", "more extensive", "more philosophical", "more maneuverable".These are phrases that describe a quality to a higher degree.Let's think about the purpose of the neuron. It seems to be recognizing these comparative/superlative structures.Possible explanations:- degree comparison words- comparative adjectives- words indicating degree- intensification and comparison- words signifying degree or comparison"words signifying degree or comparison" is 5 words."intensification and comparison" is 3 words. This is concise and captures the main theme.Let's check if "intensification and comparison" is specific enough.It's more specific than "words related to comparison", it points to the *act* of intensifying or comparing.Let's consider if there's a simpler way to phrase it.The neuron is firing when it sees things like "most X", "more Y".How about focusing on the *result* of the intensifiers?The result is a heightened or comparative state.Let's look at the TOP_POSITIVE_LOGITS again:. ; : " | padaThese don't seem to strongly align with *just* intensifiers/comparisons. However, they are often punctuation that follows phrases, or words used in specific contexts that might contain these comparisons.Let's stick with the strong pattern of "most", "more", followed by adjectives."intensifiers and comparisons" -> 3 words. Good."words of magnitude or comparison" -> 4 words.The tokens themselves: 'most', 'more', 'eless' (likely suffix for comparative/superlative), then outcomes like 'artistic', 'extensive', 'ness'. This reinforces the idea of degree.Let's also consider 'Values', 'years', 'license', 'mindset'. These are nouns.The pattern that connects 'most' and 'more' with 'artistic' and 'extensive' seems stronger than trying to find a unifying theme for *all* the `MAX_ACTIVATING_TOKENS`.The rule is: "Not ALL top activating texts/tokens have to match the exact same pattern, but a majority should."The majority of the `MAX_ACTIVATING_TOKENS` *or* the ones that show clear post-token relationships point towards degree and comparison.'most' -> 'artistic''more' -> 'extensive'Let's try: "degree or comparative language" (4 words)This feels right.Final check:- Concise (3-20 words): "degree or comparative language" is 4 words. OK.- Phrase: Yes. OK.- Specific: Captures the 'most', 'more', 'artistic', 'extensive' interaction. Yes. OK.- No forbidden phrases: "This neuron detects/predicts", "tokens", "patterns". OK.- No capitalization (unless proper noun): "degree or comparative language". OK.Could it be simpler?"comparisons and degrees" (3 words)This is also good.Let's look at the examples one last time:"most artistic", "more extensive", "more philosophical", "most artistic", "more maneuverable"These are all examples of degree or comparison.The explanation "comparisons and degrees" seems to accurately capture this. It's concise and specific to the pattern observed.Let's consider the possibility that it's indicating suffixes like '-est' or '-er'.'eless' for 'less

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/resid_post/layer_31_width_262k_l0_medium
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    েকের
    0.55
    つつ
    0.52
    蚯
    0.52
     welcomed
    0.52
    ^{\#}\
    0.48
    စျေး
    0.48
     大学
    0.47
     দিয়
    0.46
     হুইয়
    0.46
     cuore
    0.46
    POSITIVE LOGITS
    .
    0.51
     Sterile
    0.46
    :
    0.45
    ;
    0.42
     Leverage
    0.42
    "
    0.41
    RelativeTo
    0.41
     |
    0.41
    pada
    0.40
    |
    0.39
    Activations Density 0.001%

    No Known Activations