© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsinterp-engineNEWAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Gemma-3-27B
    3. 31-GEMMASCOPE-2-RES-262K
    4. 121019
    Prev
    Next
    INDEX
    Explanations

    locations for classes, events & private training.')- 'global' -> 'internal' (context: 'global internal communication')- 'as' -> 'opposed' (context: 'as opposed to')- 'dumpster' -> 'according' (context: 'dumpster according to')- 'out' -> 'In' (context: 'out In [date/year]') - this seems less strong. Let's check other instances of 'out' or similar.- 'want' -> 'to' (context: 'want to')- 'create' -> 'but' (context: 'create, but') - this is weak. Maybe the neuron detects verbs or nouns followed by connecting words?- 'famous' -> 'for' (context: 'famous for')The pattern seems to be words that are often followed by prepositions or specific connecting words that elaborate or define the preceding word.'global internal', 'as opposed', 'dumpster according to', 'want to', 'famous for'.Let's consider the TOP_POSITIVE_LOGITS: at, or, in, o, ed, et, i, u, due, id.These look like suffixes or common letter combinations. This doesn't seem directly related to the MAX_ACTIVATING_TOKENS and TOKENS_AFTER_MAX_ACTIVATING_TOKEN pattern.The primary pattern is looking at words followed by other specific words.'want to' is a strong indicator. 'famous for' is another. 'global internal' could be seen as a compound descriptor. 'as opposed' is fixed phrase.The common theme is a word followed by a word that refines it or continues a concept.Let's focus on the clear adj/noun + prep/conjunction pattern.'want to', 'famous for', 'global internal', 'as opposed'.Perhaps the neuron detects phrases or common collocations that involve prepositions or specific continuations.If we look at the first few MAX_ACTIVATING_TOKENS: '&', 'global', 'as', 'dumpster', 'out', '5', 'want', 'create', '5', 'famous'.And TOKENS_AFTER_MAX_ACTIVATING_TOKEN: 'private', 'internal', 'opposed', 'according', 'In', 'by', 'to', ',', 'by', 'for'.This is pointing towards detecting specific word pairs or common continuations.Let's try to summarize: words followed by prepositions or continuations.'want to create' and 'famous for' are examples.This neuron seems to activate when seeing certain words followed by other specific words, often indicating a relationship or continuation."want to", "famous for", "global internal"Alternative: Detects words followed by prepositions or specific conjunctions."want to", "famous for" suggests this.Let's consider the simplest interpretation of the majority.'want to' is very common. 'famous for' is also clear.'global internal' is a strong pair.Considering the list of MAX_ACTIVATING_TOKENS:'want' -> 'to''famous' -> 'for''global' -> 'internal''as' -> 'opposed''dumpster' -> 'according'These are often verb + infinitive, adjective + preposition, adjective + adjective, adverb + preposition.The pattern is words followed by specific continuations.What if we look at the *type* of word following? Many are prepositions ('to', 'for', 'by', 'according').This points to detecting words that commonly precede prepositions or specific connecting words.Let's focus on "want to" and "famous for".The phrase needs to be short."words followed by specific continuations" - too long."words followed by prepositions" - covers many, but not all (like 'internal')."common word pairings" - too general."want to" and "famous for" are very strong examples.`want to` is a common verb-infinitive structure.`famous for` is a common adjective-preposition structure.`global internal` is an adjective-adjective structure.`as opposed` is an adverb-preposition structure.The neuron detects common phrases or collocations.Perhaps "common phrases" is a good explanation.Let's check length: 2 words. It fits the criteria.However, the prompt asks for "what the neuron *detects or predicts* by finding patterns in lists."It's not just any phrase, but what it *detects*.The detected patterns are like "want to", "famous for".The output should be the *pattern itself*.If a neuron detects "want to", its behavior is like detecting "want to".The neuron seems to activate on specific word continuations."want to" is a good example."famous for" is another.Let's re-read the prompt: "concise explanation (3 to 20 words) that captures what the neuron detects or predicts by finding patterns in lists."The pattern is a word followed by a specific other word.Examples:want -> tofamous -> forglobal -> internalas -> opposedThis is like detecting common collocations or phrasal patterns.How about: "common word continuations"? 3 words.Or: "words followed by prepositions" (too restrictive).Consider the MAX ACTIVATING TOKENS: &, global, as, dumpster, out, 5, want, create, 5, famousAnd TOKENS_AFTER_MAX_ACTIVATING_TOKEN: private, internal, opposed, according, In, by, to,

    np_acts-logits-general · gemini-2.5-flash-lite

    custom command scratch

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    google/gemma-scope-2-27b-pt/resid_post/layer_31_width_262k_l0_medium
    Prompts (Dashboard)
    392,802 prompts, 256 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    튕
    0.52
     비밀
    0.48
     উৎপাদন
    0.47
     रूट
    0.46
    킴
    0.46
     народного
    0.45
    聞いた
    0.44
     уважа
    0.44
    暗
    0.44
     экрана
    0.44
    POSITIVE LOGITS
    at
    0.62
    or
    0.61
    in
    0.54
    o
    0.53
    ed
    0.51
    et
    0.51
    i
    0.51
    u
    0.48
     due
    0.48
    id
    0.48
    Activations Density 0.001%

    No Known Activations