© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 10-GEMMASCOPE-2-TRANSCODER-262K
    4. 49904
    Prev
    Next
    INDEX
    Explanations

    The phrases "Uptown Funk" and "System Uptime" frequently appear.There's a strong signal for "Uptown" followed by "Funk".There's also a signal for "ptime" followed by technical terms like "Explained" or ":".Let's try to combine these. The neuron seems to be sensitive to specific phrases or contexts."Uptown Funk" is a very strong candidate."ptime" followed by indicators of technical documentation (Expired, : **) is another.The challenge is to unify these.Maybe look at the function of the neuron.The neuron appears to be sensitive to two distinct contexts:1. Musical references, specifically "Uptown Funk".2. Technical references, specifically dealing with "uptime" in a system context.Let's try to find a unifying theme or an explanation that covers both. The user wants a *single phrase*.If I have to pick one, "Uptown Funk" is very prominent in the activating texts.The `MAX_ACTIVATING_TOKENS` shows `own` and `Uptown`.`TOKENS_AFTER_MAX_ACTIVATING_TOKEN` shows `Funk` after `own` (implicitly, as `Uptown Funk` is a phrase where `Funk` follows `Uptown`).The `ptime` tokens followed by technical terms are also present.Example: `System Uptime Explained``Website Uptime Explained`So, two main themes:A) "Uptown Funk"B) "Uptime" followed by technical descriptors (like Explained)Can I combine these elegantly?"Uptown Funk" and "Uptime Explained" is too long.Let's re-evaluate `MAX_ACTIVATING_TOKENS`:`own` (many times)`Uptown``ptime` (many times)The `own` could be part of `Uptown`.The `ptime` could be part of `Uptime`.So, key triggers are `Uptown` and `Uptime`.What follows `Uptown`? `Funk!`, `Funk`, `circles`, `event`.What follows `Uptime`? `& Stability:`, `Explained:`.The pattern seems to be a specific prefix (`Uptown`, `Uptime`) leading to a specific suffix or context.Let's look at `TOP_POSITIVE_LOGITS` again: `стью`, `propulsion`, `nings`, `架构`, `сах`, `compassing`, `ogène`, `pivoting`, `omination`, `amiseks`.This list is mostly foreign or abstract words. It's less helpful for finding a direct pattern *in English*. It suggests the neuron might have learned something general about languages or concepts that doesn't directly map to a simple English phrase derived from the activating texts, or it *uses* these abstract concepts to *signal* the specific English phrases.Given the constraints and the data:- The neuron fires strongly on words like "own", "Uptown", "ptime".- "Uptown" is followed by "Funk" or associated with events and circles.- "ptime" is part of "Uptime" and is followed by technical terms like "Explained".The most salient, distinct pattern is "Uptown Funk" (from the activating texts and `MAX_ACTIVATING_TOKENS` + `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`).The other pattern is related to "Uptime" in a technical context.Can I explain the neuron as detecting the *combination* of these two? Or perhaps it detects *one of these*.Let's focus on the two main clusters of evidence:1. `Uptown` + `Funk` (music)2. `Uptime` + Technical context (systems/web hosting)The neuron could be detecting "Uptown Funk" song mentions, or it could be detecting system "Uptime" discussions.If I *must* find a single, unifying pattern, it's difficult.However, "Uptown" appears quite often. And "Funk" follows it.And "ptime" (part of "Uptime") appears often.Let's consider the rules:- Concise (3-20 words).- Find patterns.- Not just listing tokens.- Specific.- Majority should match.The `MAX_ACTIVATING_TOKENS` has `own` many times, then `Uptown`, then `ptime` many times.The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` has `Funk` after `own`/`Uptown`."Uptown Funk" and system uptime

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_10_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    İ
    0.72
    ের
    0.70
    ों
    0.70
    II
    0.68
    ASH
    0.68
    }
    0.65
    Visit
    0.64
     Visit
    0.64
    EX
    0.63
    乜
    0.62
    POSITIVE LOGITS
    стью
    0.65
     propulsion
    0.64
    nings
    0.60
    架构
    0.57
    сах
    0.55
    compassing
    0.55
    ogène
    0.54
     pivoting
    0.53
    omination
    0.53
    amiseks
    0.52
    Activations Density 0.000%

    No Known Activations