© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-4-31B
    3. 30-RES-MATRYOSHKA-131K
    4. 116851
    Prev
    Next
    INDEX
    Explanations

    I'm going to analyze the provided lists to identify a concise pattern that describes the neuron's behavior.1. **MAX_ACTIVATING_TOKENS**: `'):', '::', ''', 'This', ';;', 'Jun'` * These tokens often appear at the beginning of a new segment, a line, or before descriptive text. `This` and `Jun` suggest starting points or references. Punctuation like `':`, `::`, `;;`, `''` seem to delineate or mark specific locations.2. **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: `'"', '(', '@', '=>', 'invention', ';;', 'present'` * `invention` and `present` are strong indicators. * `=>` suggests assignment or relation. * `"` and `(` are structural.3. **TOP_POSITIVE_LOGITS**: `which`, `コーナー`, `one`, `only`, `व्यंजनों`, `J`, `}}\`, `Bl`, `for`, `rec` * This list is quite diverse. `व्यंजनों` (recipes) and `コーナー` (corner) are specific, potentially foreign or domain-specific terms. `which`, `for`, `one`, `only` are functional.4. **TOP_ACTIVATING_TEXTS**: * `', 'George'), 'Lifeguard'` * `Lifeguard` * `"acno": "D03930", "acquisitionYear": 1856, "all_artists": "Joseph Mallord William Turner` (structured data, art attribution) * `"apiVersion": "2015-08-01", "type": "Microsoft.Web/serverfarms", "name": "[variables('site` (code/configuration) * `plan. Here is the section` * `quest_template_addon SET RequiredSkillID = 2557, RequiredSkillPoints = 100 WHERE (ID = 52227); UPDATE quest_template_addon SET RequiredSkillID = 2` (SQL/game data) * `the owner of this cycle shop.$" MauvilleCity_BikeShop_Text_181016:: @ 8181016 .string "RYDEL: Your RUNNING SHOES...\n` (game text, specific shop names/dialogue) * `c964ee0ededf28c96ebd9db5099ef910' => $vendorDir . '/` (code) * `This invention relates generally to tank bioleaching of a base metal. In tank bioleaching of a sulphide concentrate which` (patent/technical description) * `;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;` (separator) * `Jun 1, 2013 "On the whole, it is patience which makes the final difference between those` (date, quote) * `The present invention relates to data processing by digital computer, and more particularly to a user interface (UI) for defining a text` (patent/technical description)**Pattern Identification**:- The presence of "invention" and "present invention" in the text samples and `invention`, `present` in `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` is very strong.- The context often involves descriptions of new ideas, processes, or systems, which aligns with "invention".- The `MAX_ACTIVATING_TOKENS` seem to mark the start or structural elements preceding such descriptions.- The `TOP_POSITIVE_LOGITS` like `which` or `for` are functional words that would appear in such descriptive sentences. The foreign words might be noise, or representative of specialized technical domains.**Conclusion**: The neuron seems to activate for text related to descriptions of new things, especially in technical or patent contexts. "Invention" is the most prominent theme, often preceded by structural tokens or markers.**Explanation**: descriptions of invention or patents"invention" and "present invention" are key phrases. The context involves technical descriptions and patents.The shortest, most specific explanation capturing this is related to the concept of invention or new developments.Let's check the rules:- Concise (3 to 20 words): "descriptions of invention or patents" is 5 words.- Single word, phrase, or pattern: It's a phrase.- About tokens following or preceding: It describes the content the tokens point to.- Avoid listing tokens: Yes.- No "This neuron detects/predicts": Yes.- No "tokens" or "patterns": Yes.- No capitalization unless proper noun: Yes.- Specific: Yes, "invention" and "patents" are specific.- Majority match: The "invention" theme is very strong.Another option could be "technical descriptions of new things" or "patent abstract content".Given the strong presence of "invention" and "present invention", focusing on that seems most accurate.Let's consider just "invention descriptions". That's 2 words, too short."patent invention descriptions". 3 words. Specific."descriptions of invention and patents". 5 words. Specific and covers both aspects.The `TOP_POSITIVE_LOGITS` did have some unusual words but also 'which', 'one', 'for'. This suggests the neuron isn't just looking for specific words but the *context* where these words appear, and that context is heavily skewed towards presenting new ideas.The presence of `MAX_ACTIVATING_TOKENS` like `This`, `Jun`, and punctuation `':`, `::`, `;;` suggest these are often *starting points* for these descriptions.Final check:- **MAX_ACTIVATING_TOKENS**: `'):', '::', ''', 'This', ';;', 'Jun'`- **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: `'"', '(', '@', '=>', 'invention', ';;', 'present'`- **TOP_POSITIVE_LOGITS**: `which`, `one`, `for` (ignoring less clear ones)- **TOP_ACTIVATING_TEXTS**: "This invention relates...", "The present invention relates..."The most direct pattern is the "invention" related phrases found in the texts, and these are often introduced after certain markers.The phrase "invention descriptions" feels too simplistic. "descriptions of invention or patents" is good.What about just referencing the core activity? People "describe" inventions.Let's try to be more active. The neuron is finding/highlighting these."finding invention descriptions" - 3 words.The phrase "invention descriptions" is a very

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    decoderesearch/gemma-4-saes/gemma-4-31b
    Prompts (Dashboard)
    16,384 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     आपात
    -0.07
    さまざまな
    -0.07
     terutama
    -0.07
    并未
    -0.07
     மூன்று
    -0.06
     maakt
    -0.06
     क्योंकि
    -0.06
    <b>
    -0.06
     குறிப்பாக
    -0.06
    approximately
    -0.06
    POSITIVE LOGITS
     which
    0.07
    コーナー
    0.06
     one
    0.06
     only
    0.06
     व्यंजनों
    0.06
     J
    0.06
     }}\
    0.05
     Bl
    0.05
     for
    0.05
    rec
    0.05
    Activations Density 0.001%

    No Known Activations