© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-12B-IT
    3. 12-GEMMASCOPE-2-RES-16K
    4. 13721
    Prev
    Next
    INDEX
    Explanations

    - It's about legal/contractual terms and their context.- The words "clause", "provision", "provided" are strongly present in MAX_ACTIVATING_TOKENS.- The TOP_ACTIVATING_TEXTS reinforce this: "treaty provisions", "Understanding the Clause", "common clauses", "AWS EC2 Provisioning", "Provided we maintain".- The TOKENS_AFTER_MAX_ACTIVATING_TOKEN like '.', ',', 'ing', 'table' suggest these words are often followed by the end of a sentence, a clause continuation, or references to data structures like tables.The core theme is "clauses" and "provisions".Let's try to form a concise explanation.- "clauses and provisions" (3 words) - Accurate but maybe too direct.- "legal clauses and provisions" (4 words) - More specific.- "contract clauses and provisions" (4 words) - Similar specificity.- "extracting clauses and provisions" (4 words) - Implies an action from the text.- "legal clauses and provisions context" (5 words) - Adds context.The prompt asks for what the neuron *detects or predicts by finding patterns*.The neuron seems to activate highly when text contains terms like 'clause', 'provision', 'provided', which are related to legal or contractual documents.How about focusing on the *context* where these words appear?"legal and contractual provisions" is good."contractual clauses and provisions" is also good.The examples include "AWS EC2 Provisioning", which is technical but still provisioning."treaty provisions" is legal."common clauses" is general legal/contractual."policy numbers from the insurance_policy table and the policy_clause table" - this is extremely specific and connects 'clause' with 'policy' and 'table'.Given the presence of "table" in TOKENS_AFTER_MAX_ACTIVATING_TOKEN, and specific examples like "insurance_policy table and the policy_clause table", a pattern related to structured data or code might be relevant.Let's re-examine the lists with this in mind.MAX_ACTIVATING_TOKENS: provisions, Clause, provision, clauses, Provision, providedTOKENS_AFTER_MAX_ACTIVATING_TOKEN: . ** ing , from table ing . we"table" is in TOKENS_AFTER_MAX_ACTIVATING_TOKEN. This is a strong signal for data structures, databases, or code.The presence of "clause" and "provision" alongside "table" suggests something like:- "database clauses"- "policy clauses from table"- "provisioning tables"The specific examples:1. "treaty provisions" - Legal, not table-related.2. "Understanding the Clause" - Legal, not table-related.3. "On-premise often leads to over-provisioning" - Technical/operations, not table-related.4. "common clauses" - Legal/contractual, not table-related.5. "extract clauses from the following sentence" - NLP task, not table-related.6. "policy_clause table ... insurance_policy table" - EXTREMELY relevant. Connects clause, policy, table.7. "AWS EC2 Provisioning" - Technical, related to infrastructure.8. "Provided we maintain our lead" - Conditional statement, not table-related.The strongest and most specific combined signal comes from example #6 where "clause" is directly linked to "table". Also, "Provisioning" in example #7 is technical.The other tokens in TOKENS_AFTER_MAX_ACTIVATING_TOKEN like '.', ',', 'ing' are generic.But "table" stands out.What if the neuron detects specific combinations? "Clause" followed by "table"? "Provision" followed by "table"? Or maybe "clause" and "table" appearing in proximity.Let's look at TOP_POSITIVE_LOGITS again:ী, та, ن, ン, ed, ۹, uje, 마련, selffe, elif.These are mostly non-Latin characters or short sequences. This might imply the neuron detects specific token IDs that are associated with these patterns, rather than a semantic meaning of "clause" itself. For example, if token ID X corresponds to "table" and token ID Y corresponds to "clause", and Y is often followed by X, this neuron might be picking that up.The prompt's rule: "Your job is to explain the behavior of the neuron in a single short phrase. You should look at the lists and find a pattern that helps you explain the behavior of the neuron....The explanation should be specific."Let's consider "clause" and "table". The combination is present in the example texts.The prompt says "find a pattern that helps you explain the behavior".The MAX_ACTIVATING_TOKENS are predominantly "clause" and "provision" related terms.The TOKENS_AFTER_MAX_ACTIVATING_TOKEN contain "table".clause and table

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-12b-it/resid_post/layer_12_width_16k_l0_medium
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    িতে
    1.79
    它
    1.64
     flanges
    1.55
    กำ
    1.53
     CSK
    1.53
     Eigen
    1.50
    릭터
    1.50
     aisl
    1.49
    িল
    1.48
     зас
    1.48
    POSITIVE LOGITS
    ী
    1.89
    та
    1.87
    ن
    1.85
    ン
    1.80
    ed
    1.63
    ۹
    1.59
    uje
    1.54
     마련
    1.49
    zelfde
    1.48
    elif
    1.44
    Activations Density 0.055%

    No Known Activations