© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B
    3. 31-GEMMASCOPE-2-RES-262K
    4. 128336
    Prev
    Next
    INDEX
    Explanations

    letters, control, relationships, cases, eyebrows

    np_acts-logits-general · gemini-2.5-flash-lite

    listing specific examples*Self-correction during thought process:*The words `particularly` and `especially` appear in `MAX_ACTIVATING_TOKENS`.The subsequent tokens are often prepositions or phrases indicating description or scope, like `in`, `to`, `with`, `or`, `customer`, `financial`, `parameters`, `color`.The `TOP_ACTIVATING_TEXTS` frequently contain phrases like "particularly in areas related to...", "list types of information - e.g., business plans, customer lists, financial information...", "Trying to self-treat...", "Same parameters as first call...", "especially with Dyonasian aspect...".The core pattern seems to be introducing a specific detail or list of items/examples after a modifying word or a verb indicating effort/action.Let's refine:- `particularly` + `in areas related to X, Y, Z.`- `Trying` + `to do something.`- `[list types of information - e.g., A, B, C]`- `sleeves` (context for item description)- `Same parameters`This neuron seems to be about specifying details or providing particular examples.The `TOP_POSITIVE_LOGITS` (`Hilo`, `Sac`, `Aqu`, `Humidity`, `Staff`, `Colombo`, `Arquivo`, `ACO`, `Salon`, `rosters`) are not strongly related to this pattern of listing specifics. They are more like names or labels.Considering `MAX_ACTIVATING_TOKENS`: `particularly`, `Trying`, `Same`, `sleeves`, `especially`.Considering `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`: `in`, `customer`, `to`, `parameters`, `financial`, `color`, `or`, `with`.Considering `TOP_ACTIVATING_TEXTS`: many examples involve listing specific types of information, physical attributes, or actions. listing specific examples

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-pt/resid_post/layer_31_width_262k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    ซึ่ง
    0.49
     März
    0.48
    ศาส
    0.45
    
    0.43
    を使用
    0.42
     który
    0.41
    (".
    0.41
     yılında
    0.41
    arın
    0.41
     যৌ
    0.40
    POSITIVE LOGITS
    Hilo
    0.51
     Toledo
    0.46
    Staff
    0.46
    Sac
    0.46
    Salon
    0.46
    Checkbox
    0.46
    Furn
    0.46
     Cots
    0.46
    Fols
    0.46
    Panorama
    0.45
    Activations Density 0.001%

    No Known Activations