© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 13-GEMMASCOPE-2-TRANSCODER-262K
    4. 164046
    Prev
    Next
    INDEX
    Explanations

    Based on the provided lists, the neuron appears to be strongly associated with detecting the word "Koin" and related terms, especially in technical contexts like dependency injection or programming libraries, and also terms related to "Rate" or percentages.Let's break it down:* **MAX_ACTIVATING_TOKENS**: contains `oin`, `k`, `Rate`, `rates`, `rate`. "oin" is very similar to "Koin". "Rate" is also present.* **TOKENS_AFTER_MAX_ACTIVATING_TOKEN**: contains `.`, `:**`, `in`. This suggests sequences like "Koin.", "Koin:**", "Koin in".* **TOP_POSITIVE_LOGITS**: contains `Koin`, `Service`, `User`, `Language`. "Koin" is a specific library mentioned in the text. Other terms are more general.* **TOP_ACTIVATING_TEXTS**: explicitly mentions `Koin` in the context of `Dependency Injection: Hilt (Recommended) or Koin`. It also frequently mentions `Rate` (e.g., `Penetration Rate`, `mobile penetration rate`).The most consistent and specific pattern points to "Koin" and "Rate".Considering the rules:- Conciseness (3-20 words)- Find patterns, not just list tokens- No

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/transcoder_all/layer_13_width_262k_l0_small_affine
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     bustling
    0.61
     walks
    0.59
     walkers
    0.59
     aandacht
    0.59
     jogging
    0.58
     Tanaka
    0.58
     kittens
    0.58
     što
    0.57
     walker
    0.57
     contacto
    0.57
    POSITIVE LOGITS
    New
    0.50
    Airplane
    0.50
    User
    0.50
    Service
    0.50
    Language
    0.49
    1
    0.49
    Soil
    0.48
    MetaData
    0.48
    Woman
    0.48
    Ht
    0.48
    Activations Density 0.000%

    No Known Activations