© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Qwen3-32B
    3. 32-RESID-BATCHTOPK-65K
    4. 3585
    Prev
    Next
    INDEX
    Explanations

    of a modest nature"- "...social nature"- "...of an increasing importance..."- "...of mesenchymal origin"- "...of industrial origin"- "...of Hebrew origin"- "...of high grade"- "...of another origin"- "...or cultural in nature"The phrase "X of Y nature" or "X of Y origin" seems to be a strong pattern.'nature' and 'origin' are prominent in the `TOP_POSITIVE_LOGITS`.In `MAX_ACTIVATING_TOKENS`, 'origin', 'of', 'in', 'an', 'a' are frequent.In `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`, 'nature', 'increasing', 'high', 'another', 'modest' appear.Combining these, the neuron seems to identify phrases that describe the quality or source of something, often in the "X of Y nature" or "X of Y origin" structure.Let's try to capture this concisely."of X origin" or "of X nature" is the core.Possible explanations:- X nature or X origin- descriptions of nature or origin- quality or type of something- specifies nature or origin- nature of or origin ofLooking at the `TOP_POSITIVE_LOGITS`, 'nature' is very strong. 'origin' is strong in `MAX_ACTIVATING_TOKENS`. The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` also shows descriptive words like 'modest', 'increasing', 'high', 'another'.The pattern appears to be identifying phrases that describe the essence, source, or type of something."X of Y nature" or "X of Y origin"."nature" and "origin" are key.Consider the structure: `[Article/Preposition] [Description] [nature/origin]`e.g., "of a modest nature", "of industrial origin"The neuron seems to be capturing the descriptive part following "of a/an" or just "of".Let's re-examine the core elements:- `MAX_ACTIVATING_TOKENS`: 'origin', 'of', 'in', 'an'- `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`: 'nature', 'increasing', 'high', 'another', 'modest'- `TOP_POSITIVE_LOGITS`: 'nature', 'Nature'The common theme is describing the *kind* or *source* of something."describing nature or origin" - 4 words."nature of or origin of" - 5 words."kind or source described" - 4 words."specifies nature or origin" - 4 words.The provided examples focus on "nature" and "origin" as positive logits and `MAX_ACTIVATING_TOKENS`.The tokens after are descriptors for nature/origin."nature or origin descriptors" - 4 words. This seems quite fitting.Let's check the constraints. 3-20 words. Not start with "This neuron detects/predicts". No "tokens" or "patterns". Specific."nature or origin descriptors" fits these.Alternative:"qualities of nature or origin" (5 words)"describes nature or origin" (4 words)Let's pick the most direct one reflecting the frequent terms."nature or origin descriptors" directly uses the identified key terms and their function.The prompt asks for what the neuron *detects or predicts by finding patterns in lists*. The lists suggest the neuron is sensitive to the *words* "nature" and "origin" and whatever *follows* them.Let's look at `TOP_POSITIVE_LOGITS` again: `æĢ§è´¨`, `nature`, `nature`, `Nature`, `Nature`, `ä»ħä¾Ľ`, `Mercer`, `ä¼ļè§ģ`, `type`, `æĻ®æĥł`.'nature' and 'type' are here. 'æĢ§è´¨' might be 'quality' in Chinese. 'ä»ħä¾Ľ' might be 'origin' in Chinese.This confirms the focus on quality, type, nature, origin."quality, type, nature, or origin" would be too broad and long. The prompt asks for *finding patterns*.The pattern is `[descriptor] nature` or `[descriptor] origin`.The neuron seems to highlight words that are modifiers before "nature" or "origin", or the words "nature" and "origin" themselves.Consider the structure `of [descriptor] nature`:"of a modest nature""social nature""of an increasing importance" (similar structure)"of mesenchymal origin""of industrial origin""of Hebrew origin""of high grade" (similar structure)"of another origin"The neuron is detecting the structure `of [X] Y` where Y is related to nature/origin/type/quality, and X is a descriptor."describing quality or source" - 4 words."nature or origin descriptions" - 4 words."type, nature, or origin descriptors" - 6 words.The examples point heavily to "nature" and "origin".Let's go with a simpler phrasing that captures the essence."nature, origin, or type" - 4 words.This covers the positive logits and the themes.The actual examples show phrases like "of a X nature", "of Y origin".The neuron is strongly associated with finding the *words* 'nature' and 'origin' and the descriptive words that often precede them or are associated with them."nature and origin descriptions" seems good."nature or origin descriptors" seems also good.Let's stick to describing *what* it activates on.What about "X of Y type"? Look at 'type' in logits.The `MAX_ACTIVATING_TOKENS` `a, a, in, an, origin, origin, origin, of, of, in` suggests prepositions and words like 'origin'.`TOKENS_AFTER_MAX_ACTIVATING_TOKEN` `modest, hy, nature, increasing, ;, \, ., high, another, nature` suggests the descriptive words.So, the pattern is `[preposition/article] [descriptor] [nature/origin/type]`.The neuron identifies these descriptive phrases."descriptors of nature or origin" - 5 words."nature or origin qualifiers" - 4 words.Refining based on the structure seen in `TOP_ACTIVATING_TEXTS`:- "modest nature"- "social nature"- "increasing importance"- "mesenchymal origin"- "industrial origin"- "Hebrew origin"- "high grade"- "another origin"The neuron seems to identify phrases that specify the *kind* or *source* of something."kind or source descriptions" - 4 words.Let's try to be even more specific to the found words:"nature or origin phrases" - 4 words.This is extremely direct.Final check:- Concise (3 to 20 words): "nature or origin phrases" is 4 words. OK.- Phrase, not a full sentence: OK.- Specific: Yes, refers

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    adamkarvonen/qwen3-32b-saes/saes_Qwen_Qwen3-32B_batch_top_k/resid_post_layer_32
    Prompts (Dashboard)
    16,384 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    罩
    -0.09
    èĩ§
    -0.09
    æĬĺ
    -0.08
    æ´»å¾Ĺ
    -0.08
    骤
    -0.08
    ек
    -0.08
    noinspection
    -0.08
    åŁ¹èĤ²
    -0.08
    没æľīæĥ³åΰ
    -0.08
    rust
    -0.08
    POSITIVE LOGITS
    æĢ§è´¨
    0.19
     nature
    0.18
    nature
    0.16
    Nature
    0.13
     Nature
    0.12
    ä»ħä¾Ľ
    0.10
     Mercer
    0.10
    ä¼ļè§ģ
    0.09
     type
    0.09
    æĻ®æĥł
    0.09
    Activations Density 0.141%

    No Known Activations