© Neuronpedia 2026
    Privacy & TermsBlogGitHubSlackTwitterContact
    Neuronpedia logo - a computer chip with a rounded viewfinder border around it

    Neuronpedia

    Jacobian LensNEW
    Natural Language
    Autoencoders
    NEW
    Assistant AxisNEWCircuit TracerUPDATESteerSAE EvalsExportsAPI Community BlogPrivacy & TermsContact
    1. Home
    2. Olmo-3-1025-7B
    3. 16-RES-MATRYOSHKA-65K
    4. 55966
    Prev
    Next
    INDEX
    Explanations

    meme-selfandtellsandimpyellinglevelbrainThis list primarily contains pronouns like "me" and "-self", conjunctions like "and", and words related to communication or internal states like "tells", "yelling", "brain". There's a strong self-referential ("me", "-self") and expressive ("tells", "yelling", "brain") component.**TOKENS_AFTER_MAX_ACTIVATING_TOKEN:**me thoughtme ,-self ,and Itells meand Iimp canyelling atlevel itbrain hasThe tokens following the MAX_ACTIVATING_TOKENS often complete phrases related to personal thoughts ("me thought"), personal statements ("tells me", "and I"), or actions/states ("yelling at", "brain has").**TOP_POSITIVE_LOGITS:**saysSaytellssayingbeginsuggestsaku (Japanese for 'I')crysuggestolf (less clear)This list strongly reinforces the theme of communication, expression, and opinion. Words like "says", "tells", "suggests", and "cry" point to a neuron that activates when someone is speaking, expressing, or indicating a state. The inclusion of "aku" (Japanese for 'I') further emphasizes the personal or self-referential aspect.**TOP_ACTIVATING_TEXTS:**- "part of me thought he was just talking bullshit"- "23-year old me, a few things..."- "my online-shopping-self, it’s kind of a ‘one step forward,"- "it bothers me."- "Experience tells me that people don’t change jobs solely for money"- "what I thought I could do myself."- "own happiness" (related to self)- "yelling at me"- "don’t align with..." (communication/opinion context)- "my brain has been regurgitating"The texts confirm the self-referential aspect ("me", "myself", "my online-shopping-self") and the theme of expressing thoughts, opinions, or internal states ("thought", "tells me", "yelling at me", "brain has been regurgitating").**Pattern Identification:**The neuron seems to activate when the text describes personal thoughts, introspections, or expressions, often related to the self ("me", "-self", "myself"). The subsequent tokens and positive logits frequently point towards making statements, having thoughts, or conveying information ("thought", "tells", "says", "suggests", "I").**Synthesis:**The neuron appears to be sensitive to instances where the speaker or writer is expressing personal thoughts, opinions, or reflections, often directed towards themselves or stating something directly.Final explanation should be 3-20 words and a phrase.Potential phrases:- expressing personal thoughts and feelings- self-reflection and direct statements- personal narration and opinion- expressing "me" and future actions- what "I" think or say"self-reflection and direct statements" seems to capture both the internal aspect ("me thought", "brain regurgitating") and the output aspect ("tells", "says", "suggests").Let's refine to be more concise and direct, focusing on the core action."personal thoughts and statements" (4 words)"self statements and reflections" (4 words)"expressing personal thought or statement" (5 words)The neuron is picking up on the act of stating one's own thoughts or feelings."personal statements and introspection" is good."self-statements and thoughts""personalized thoughts and expressions""me saying or thinking"Looking at "TOP_POSITIVE_LOGITS": says, Tells, suggests, cry. These are all active expressions.Looking at "MAX_ACTIVATING_TOKENS": me, -self. This is the subject.Looking at "TOKENS_AFTER_MAX_ACTIVATING_TOKEN": thought, I, me. These are often the immediate completion.Combining: "me" + "thought" or "me" + "tells" or "self" + something.The neuron is concerned with the self being the origin of statements or thoughts."self statements and personal thought" (5 words)"personal thoughts and sayings" (4 words)"self-referential statements and thoughts" (4 words)"self statements and thoughts" is very direct and covers the key elements.Let's check the word count and rules.- Concise (3 to 20 words): "self statements and thoughts" is 4 words. Yes.- Phrase, not a sentence: Yes.- Captures what neuron detects/predicts by finding patterns: Yes, identifies self-referential statements/thoughts.- Avoid listing tokens: Yes.- Avoid "tokens" or "patterns": Yes.- No capitalization unless proper noun: Yes.- Specific: Yes.Consider "personal thoughts and statements". This is also 4 words. The examples show things like "part of me thought", "tells me that", "yelling at me", "my brain has been regurgitating". These are all personal experiences being communicated."personal statements and introspection" seems to fit well."messages from the self" - too poetic."self communication and expression" - 4 words. This feels strong. It covers "me" and "tells/says/suggests".Let's look at TOP_POSITIVE_LOGITS again: says, Say, tells, saying, suggest, suggest. These are about *saying* or *suggesting*.MAX_ACTIVATING_TOKENS: me, -self. This is the *self*.TOKENS_AFTER: thought, I, me. Completes the thought/statement.So the neuron is about the *self* communicating or expressing *thoughts*."self communication and expression" seems accurate.Another pass:The neuron activates on "me" or "self" followed by expressions of thought, telling, or saying.Examples: "part of me thought", "23-year old me, a few things you don't want really want to hear from the 33-year old me", "bothers me", "tells me that", "yelling at me", "my brain has been regurgitating".These are all about the self talking, thinking, feeling, or having something happen to it that is being related."personal experience and expression" is 4 words."self-directed thought and speech" is 5 words."personal statements and reflections" is 4 words.I like "personal statements and reflections" or "self-communication and expression.""Self-communication and expression" seems to capture the "me" / "-self" with "says" / "tells" / "yelling" very precisely.Let's re-evaluate against the prompt "find patterns in lists".MAX_ACTIVATING_TOKENS: me, -self (self-reference)TOKENS_AFTER: thought, me, I (completion of thought/statement)TOP_POSITIVE_LOGITS: says, tells, suggest, cry (expression/communication)The pattern is: the self is expressing something."self statements or expressions" (4 words)

    np_acts-logits-general · gemini-2.5-flash-lite
    New Auto-Interp
    Top Features by Cosine Similarity
    Configuration
    decoderesearch/olmo-3-saes/olmo-3-1025-7b
    Prompts (Dashboard)
    24,576 prompts, 128 tokens each
    Dataset (Dashboard)
    monology/pile-uncopyrighted
    No Configuration Found
    Embeds
    IFrame
    Link
    Not in Any Lists

    No Comments

    Negative Logits
    imits
    -0.09
    .TXT
    -0.08
     piv
    -0.08
    [o
    -0.08
     fir
    -0.08
     Ends
    -0.08
    (...)
    -0.08
     footing
    -0.07
     Liberation
    -0.07
     Lines
    -0.07
    POSITIVE LOGITS
     says
    0.10
     Say
    0.09
     tells
    0.09
     saying
    0.09
    egin
    0.09
     suggests
    0.09
    aku
    0.09
    cri
    0.09
     suggest
    0.08
    olf
    0.08
    Activations Density 0.036%

    No Known Activations