© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-4-31B
    3. 30-RES-MATRYOSHKA-131K
    4. 118218
    Prev
    Next
    INDEX
    Explanations

    company names and technical terms (e.g., `App`, `Cargo`, `Cascade`, `Image`, `Spatial`) are very prominent in `MAX_ACTIVATING_TOKENS`.The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` often seem to be continuations of the preceding token, sometimes forming compound words or phrases:- App -> Service- var -> ints- Cargo -> Corp (though "Cargo" is followed by "CAO airline designator" in text)- Cascade -> Corp- Image -> Mag (as in Image Magick)- Spatial -> analytics**TOP_POSITIVE_LOGITS**: Many are foreign words or words related to neutral/positive descriptors. "also" is present.**TOP_ACTIVATING_TEXTS**:- `AppServiceLayout`- `varints`- `Cargo` -> `CAO airline designator`- `Cascade Corp.`- `Image Magick`- `Blank, Yoga`- `Rodriguez`- `Tomorrow`- `Spatial analytics`**Concise Explanation Derivation**:The neuron seems to heavily activate on words that are often followed by specific continuations, frequently forming common phrases or compound words related to company names, technical terms, or concepts.* "App" -> "Service"* "var" -> "ints"* "Cargo" -> (continues with context)* "Cascade" -> "Corp"* "Image" -> "Mag"* "Spatial" -> "analytics"The `TOP_POSITIVE_LOGITS` are difficult to integrate directly into a simple pattern, but the consistent observation across `MAX_ACTIVATING_TOKENS`, `TOKENS_AFTER_MAX_ACTIVATING_TOKEN`, and `TOP_ACTIVATING_TEXTS` points towards sequences that form recognized terms or phrases.Let's try to capture this. The *first* token in `MAX_ACTIVATING_TOKENS` is "App". The *first* token in `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` is "Service". The text shows `AppServiceLayout`. This suggests a pattern of starting words followed by specific continuations.Consider "App", "var", "Cargo", "Chromat" (as in Chromatotherapy), "Cascade", "Image", "Blank", "Rodriguez", "Tomorrow", "Spatial".Followed by "Service", "ints", ",", "otherapy", "Corp", "Mag", ",", ")", "'", "analytics".The pattern is that these starting tokens are often followed by specific words that complete them into common terms or proper nouns.Examples:- App -> Service (`AppServiceLayout`)- var -> ints (`varints`)- Cascade -> Corp (`Cascade Corp.`)- Image -> Mag (`Image Magick`)- Spatial -> analytics (`Spatial analytics`)The explanation should capture this "word followed by its common continuation" idea.Possible explanations:- word followed by common continuation- recognized word sequences- words initiating phrases- word + suffix completion- term + specific suffix- compound word formation"recognized word sequences" or "words initiating phrases" seem good. "term + specific suffix" is also very accurate.Let's check the word count."recognized word sequences" - 3 words. Fits."words initiating phrases" - 3 words. Fits."term + specific suffix" - 3 words. Fits.Looking at top logits, "also", "prudent", "meticulous" seem like general positive descriptors, potentially indicating that the completed phrases are often neutral or positively perceived. The foreign words suggest generalization across languages.Let's refine based on the specific examples:"App" -> "Service" (App Service)"Cascade" -> "Corp" (Cascade Corp)"Image" -> "Mag" (Image Magick)"Spatial" -> "analytics" (Spatial Analytics)These are specific pairings. "term + specific suffix" captures this well.Let's consider other possibilities. "App" -> "Service", "var" -> "ints".The general idea is a word that strongly predicts what comes next, forming a compound or a standard phrase."predicting word continuations" - 3 words."word + completion" - 2 words. Too short, maybe."predicting common phrases" - 3 words.The neuron finds a word and predicts its specific, often compound, continuation.Let's focus on the concrete examples of *pairing*.App + ServiceCascade + CorpImage + MagSpatial + analyticsThis is about specific words being followed by specific other words to form something recognizable.How about "specific word pairings"? Too short."predicting specific word pairings" - 3 words."identifying common word pairings" - 4 words.Let's re-examine the prompt: "finds patterns in lists." "what the neuron detects or predicts by finding patterns in lists."The `MAX_ACTIVATING_TOKENS` are *starting points*. The `TOKENS_AFTER_MAX_ACTIVATING_TOKEN` are *potential continuations*. The `TOP_ACTIVATING_TEXTS` show *realizations* of these continuations.The pattern is: 'word' -> 'specific follow-up word'.This looks like predicting compound words or standard abbreviations/terms.* App -> Service* Cascade -> Corp* Image -> Mag* Spatial -> analyticsThe phrases like "term + specific suffix" or "predicting common phrases" seem to capture this.Let's look at the foreign words in `TOP_POSITIVE_LOGITS`. They are often adverbs or adjectives (also, prudent, meticulous). This might suggest the *context* where these pairings appear is often descriptive or informative, but the core function seems to be the *pairing* or *completion*."predicting common word sequences" - 4 words. Captures the essence of `word` + `continuation`.The foreign words are a bit of a distraction if the main pattern is linguistic completion.What if I focus on the *type* of words that are paired? Often they are proper nouns (company names, software names) or technical terms."company and technical term completion" - 5 words."identifying term completions" - 3 words.Let's test "identifying term completions"."App" -> "Service" -> App Service (term completion)"Cascade" -> "Corp" -> Cascade Corp (term completion)"Image" -> "Mag" -> Image Magick (term completion)"Spatial" -> "analytics" -> Spatial analytics (term completion)This seems to fit very well. It's concise and specific.It implies finding a term (`MAX_ACTIVATING_TOKENS`) and predicting its completion (`TOKENS_AFTER_MAX_ACTIVATING_TOKEN`).Final check on rules:- Concise (3 to 20 words): "identifying term completions" is 3 words. OK.- Captures what neuron detects/predicts: Yes, it detects terms and predicts their common completions.- Finds patterns: Yes, the pattern of term + specific suffix.- No extra phrases like "This neuron detects". OK.- No "tokens" or "patterns". OK.- Not capitalized unless proper noun. OK.- Specific. "identifying term completions" is specific. OK.- Majority should match. The strong examples cover majority. OK.- If cannot guess, return first

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    decoderesearch/gemma-4-saes/gemma-4-31b
    Prompts (Dashboard)
    16,384 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    Argument
    -0.06
    ˶
    -0.06
    จะ
    -0.06
     subordination
    -0.06
     França
    -0.05
     Aj
    -0.05
    ks
    -0.05
     bấm
    -0.05
     éché
    -0.05
     ks
    -0.05
    POSITIVE LOGITS
     консу
    0.07
     også
    0.06
     meticulous
    0.06
     deres
    0.06
    但也
    0.06
     تنها
    0.06
     также
    0.06
     also
    0.06
     prudent
    0.06
     aceasta
    0.06
    Activations Density 0.052%

    No Known Activations