© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B-IT
    3. 31-GEMMASCOPE-2-RES-65K
    4. 46685
    Prev
    Next
    INDEX
    Explanations

    flour. dough. Let's analyze the provided lists to find a pattern.* **`MAX_ACTIVATING_TOKENS`**: Contains `flour` multiple times, `for`, `teachings`, `grows`.* **`TOKENS_AFTER_MAX_ACTIVATING_TOKEN`**: `flour` is often followed by `.`. `teachings` is followed by `of`. `grows` is followed by `for`. `for` is followed by `God`.* **`TOP_ACTIVATING_TEXTS`**: Mentions "flour", "yeast mixture", "dough", "creation", "hair grows", "anagen phase". It seems to cover cooking/baking instructions and biological growth.* **`TOP_POSITIVE_LOGITS`**: Contains punctuation and characters from different languages, suggesting it might be a very general pattern or related to formatting/structure, but less helpful for a semantic explanation.Observing the `MAX_ACTIVATING_TOKENS` and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`, we see `flour` followed by a punctuation mark like `.`. This suggests the end of a step or instruction, especially in contexts like recipes. The texts also mention ingredients and steps.A strong pattern is `flour.` which appears in cooking contexts, often marking the end of a step or phrase.Let's consider explanations based on this:- "flour." is a literal sequence, but the explanation should go beyond just listing tokens and find a semantic meaning if possible.- The presence of `flour` itself is a key indicator.- The texts also touch upon `creation`, `God`, `teachings`, and `grows`, suggesting topics related to life, belief, or natural processes.However, the most consistent and direct pattern from the token lists is `flour.` and general mentions of `flour` in contexts of preparation or description.Given the examples, particularly the repetition of "flour" in `MAX_ACTIVATING_TOKENS` and its typical conclusion with a period in cooking contexts from `TOP_ACTIVATING_TEXTS`, a concise explanation relating to ingredients and steps seems appropriate.Let's look at the explicit instructions: "finding patterns in lists", "concise explanation (3 to 20 words)", "find a pattern that helps you explain the behavior", "about tokens following or preceding certain tokens", "about words starting with a sequence", "just say the pattern itself".The pattern `flour.` is very specific. The texts also show `flour` as a key ingredient.Consider the specific pairings:- `flour` followed by `.` (likely end of a step in recipes)- `teachings` followed by `of` (context for teachings)- `grows` followed by `for` (context for growth duration)The most frequent and concrete item seems to be `flour`. In the context of the given texts, `flour` is often part of a recipe instruction, and the period signifies the end of that instruction.Therefore, a pattern like "flour." captures a specific sequence observed and implies a contextual meaning within the type of texts provided.Let's re-evaluate. The TOP_POSITIVE_LOGITS are very diverse, but they might indicate what *kind* of information this neuron is sensitive to activating on when the foreground patterns are present. These lists include punctuation (`;`, `。`), non-latin scripts, and words that are grammatically function-neutral or very basic.The `MAX_ACTIVATING_TOKENS` has `flour` repeated. `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` shows `flour` followed by `.`.The `TOP_ACTIVATING_TEXTS` includes recipes (flour, dough) and religious/biological text (teachings, God, hair grows).The prompt asks for what the neuron *detects or predicts by finding patterns in lists*.Focusing on `MAX_ACTIVATING_TOKENS` and `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`:- `flour .`- `flour .`- `flour .`This specific sequence is very strong. The surrounding texts confirm `flour` is a significant word and often appears in instructional formats ending with punctuation.Let's try to generalize from `flour.` to a broader concept if possible, but the instruction also says "The explanation should be specific."If we look at the *types* of sequences:- Ingredient + punctuation (`flour .`)- Item + conjunction (`teachings of`)- Verb + preposition (`grows for`)The most prominent recurring element is `flour`. And its common successor is `.`.So, "flour." is a very direct observation.Could it be related to listing items before a concluding punctuation?"flour." appears 3 times."teachings of""grows for"The strongest, most literal pattern is "flour.", but is that too simple? The request is to find a pattern.Let's consider the other recurring elements in `MAX_ACTIVATING_TOKENS`: `for`, `teachings`, `grows`. These are not as frequent as `flour`.What if the neuron is sensitive to sequences that describe actions or states, often followed by punctuation or conjunctions?"flour." fits this."teachings of" fits this."grows for" fits this.Consider the `TOP_POSITIVE_LOGITS`: punctuation such as `;` and `。` are prominent. This reinforces the idea that the neuron might be sensitive to structures ending in punctuation or specific grammatical constructions.Therefore, "flour followed by punctuation" is a pattern.Or more concisely, combining the specific word and its common successor: "flour

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-it/resid_post/layer_31_width_65k_l0_medium
    Prompts (Dashboard)
    238,145 prompts, 512 tokens each
    Dataset (Dashboard)
    lmsys + oasst1
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
     misuse
    0.52
     misused
    0.49
    serviceTime
    0.48
     abuses
    0.46
    忧
    0.45
    肃
    0.45
    ಂತಹ
    0.45
     democracies
    0.45
     potholes
    0.44
     willingly
    0.44
    POSITIVE LOGITS
    ;
    0.48
    。
    0.48
     چھ
    0.47
    igli
    0.47
    を作成
    0.47
    श्रे
    0.46
    ロ
    0.46
     maggior
    0.46
     நடைபெற்றது
    0.46
    github
    0.45
    Activations Density 0.008%

    No Known Activations