© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 13-GEMMASCOPE-2-TRANSCODER-262K
    4. 45131
    Prev
    Next
    INDEX
    Explanations

    `IF` and `AS` seem to stand out as potential indicators of conditions or stipulations.4.`TOP_ACTIVATING_TEXTS`: * "consider ASIL decomposition. This means dividing the requirement into parts with lower ASILs, provided the overall safety goal is still met." - Mentions "provided" and "as long as". * "As long as it's not about foxes... or vacuum cleaners..." - Mentions "As long as". * "As long as it's a private, consensual activity that doesn't harm anyone, they see it as his personal business." - Mentions "As long as". * "as long as you do so in a safe and responsible manner." - Mentions "as long as". * "as long as you do not stop." - Mentions "as long as". * "as long as there is no touching other places on her body, no kissing and it has to be emotionless sex." - Mentions "as long as". * "provided you have the base LLaMA weights." - Mentions "provided". * "as long as it does not involve SCP-055." - Mentions "as long as". * "as long as you're not selling them or using them to promote a business." - Mentions "as long as". * "as long as you continue to meet the visa conditions." - Mentions "as long as".The overwhelming pattern is the phrase "as long as", indicating a condition or stipulation. The `TOP_POSITIVE_LOGITS` also might hint at conditional language (`IF`, `AS`). The `MAX_ACTIVATING_TOKENS` and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` contain words like `not` and `stop`, which often appear within these conditional statements (e.g., "as long as you do *not* stop", "as long as it does *not* involve").Therefore, the neuron seems to be related to conditional statements or requirements.as long as

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_13_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     हमारा
    0.56
    ၍
    0.54
     உற்பத்தி
    0.54
     фаразлау
    0.51
    繁殖
    0.51
     preparazione
    0.50
     செறிவு
    0.50
     Mesa
    0.50
     জাহাঙ্গীর
    0.50
     tafel
    0.49
    POSITIVE LOGITS
    IF
    0.59
    AC
    0.54
    ın
    0.54
    AS
    0.50
    ML
    0.50
    𝒾
    0.49
    IN
    0.48
    Orth
    0.48
    埜
    0.48
    İ
    0.48
    Activations Density 0.017%

    No Known Activations