© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B
    3. 31-GEMMASCOPE-2-RES-262K
    4. 122512
    Prev
    Next
    INDEX
    Explanations

    claims. The neuron appears to be detecting or predicting phrases related to legal claims or objections, especially in contexts of data processing, legal rights, or financial transactions.Let's break it down:1. **MAX_ACTIVATING_TOKENS**: `claims`, `reasons` (related to objections/claims)2. **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: `.` (often follows phrases ending sentences), `related` (used with 'reasons' in the texts)The `TOP_POSITIVE_LOGITS` lists words like "against", "fought", which could be loosely related to disputes or claims, but the strongest signal comes from the `MAX_ACTIVATING_TOKENS` and how they appear in context.Looking at specific texts:- "...exercise or defence of **claims**."- "...for **reasons** related to your personal situation..."- "...for **reasons** related to their particular situation..."The token "claims" from MAX_ACTIVATING_TOKENS is a strong indicator. The token "reasons" also appears, and in the text, it's often followed by "related to", which is in TOKENS_AFTER_MAX_ACTIVATING_TOKEN.Therefore, a good explanation would focus on "claims".If I need to be more specific based on the *combination* of tokens and context, I see things like "defence of claims" or "exercise or defence of claims". However, the request is for a *concise explanation* of what the neuron detects or predicts.Let's re-evaluate `MAX_ACTIVATING_TOKENS`: "claims".Let's re-evaluate `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`: "related", ".".The word "claims" stands out prominently in `MAX_ACTIVATING_TOKENS`. From the `TOP_ACTIVATING_TEXTS`, we see explicit mentions of "claims" in a legal/financial context.Consider the rules:- Concise (3 to 20 words).- Find a pattern.- Specific.- Do not start with "This neuron detects/predicts".- Do not mention "tokens" or "patterns".The most direct and specific pattern is the concept of 'claims'.Let's check if "claims" alone is specific enough. Given the texts, it refers to legal/financial claims.Let's look for other strong patterns."of" is very common. "3" is an outlier."payment" appears in the texts, and "pay" is in MAX_ACTIVATING_TOKENS. "comfortably" and "transmission" also appear near "pay".However, "claims" seems more central.Let's try to combine: "defence of claims" or "exercise of claims". That's a bit long.What about just the core concept?The prompt gives `MAX_ACTIVATING_TOKENS`:ofthoseends.payduringreasonsclaims3inAnd `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`:thethighsallowCertainlycomfortablytransmissionrelated.betweentheThe word `claims` is itself a strong candidate.The word `reasons` is also present.In `TOP_ACTIVATING_TEXTS`, we see:- "...defence of claims."- "...for reasons related to..."- "...details of payment."- "...during transmission..."- "...pay comfortably..."- "...chamfered ends allow..."The word `claims` is quite direct.The combination of `reasons` and `related` is also strong.However, looking at the *most prominent* items often helps. `claims` is a direct word.In the TOP_POSITIVE_LOGITS, we have "against", "fought", which could relate to disputes or claims.Let's try to focus on the strongest single concept. "claims" is very specific.If I consider `claims` and `reasons related to`, it becomes phrases like "claims" or "reasons related to".Let's look at what the prompt wants me to output: "just say the pattern itself, and do not start with phrases like..."The core concept is legal claims and potentially objections ("reasons")."claims" is a single word. Is it specific enough? Yes, in the context of the provided texts.What about combining "claims" and "objections"? The neuron might be about disputes or arguments.`TOP_POSITIVE_LOGITS` has `against`, `fought`.Let's reconsider `MAX_ACTIVATING_TOKENS`: `claims`, `reasons`.`TOKENS_AFTER_MAX_ACTIVATING_TOKEN`: `related`, `.`The phrase "defence of claims" appears in one text."Right to object: ... for reasons related to..."This neuron seems to be related to legal assertions and objections.Perhaps "claims or objections" is the best conceptual fit. However, length is important."claims" is 1 word."reasons related" is 2 words."claims and reasons" is 3 words.Let's look again at the TOP_POSITIVE_LOGITS.against鬥yalمتgiderfoughtkämpSplashconstructedajalThese are very diverse, many are non-English. This makes it harder to find a single linguistic pattern from them. The real strength must come from the MAX/TOKENS lists.The most direct and specific term is `claims`

    np_acts-logits-general · gemini-2.5-flash-lite

    app store and data structures

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-pt/resid_post/layer_31_width_262k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     econometric
    0.47
     optomet
    0.46
    unstyled
    0.46
    erView
    0.46
     কাজেই
    0.44
    、《
    0.43
    ClearedBy
    0.43
    theoretic
    0.42
     BLM
    0.42
     عصر
    0.41
    POSITIVE LOGITS
     against
    0.51
    鬥
    0.50
    yal
    0.50
    مت
    0.49
     gider
    0.48
     fought
    0.47
     kämp
    0.46
    Splash
    0.46
     constructed
    0.45
     ajal
    0.45
    Activations Density 0.001%

    No Known Activations