OpenAI's Automated Interpretability from paper "Language models can explain neurons in language models". Modified by Johnny Lin to add new models/context windows.
Default prompts from the main branch, strategy TokenActivationPair. Uses top 10 deduplicated activations.
Recent Explanations
Sentences and phrases about personal growth, therapy, healing, and emotional support—self-help / counselling language focused on feelings, confidence, choices, and recovery.
gpt-5-mini
voice so that you can make clear decisions?↵↵Welcome!
words that signal strong subjective emphasis or novelty (emotive/adjective or adverbial hype like "new", "magically", "never-before-seen", or other high-intensity descriptors).
gpt-5-mini
collection. It will not magically add brand new, never
The neuron detects explanatory/encyclopedic phrasing that describes what something is used for or how it functions (phrases about use, purpose, interaction, or tasks).
gpt-5-mini
interact with one another and with other systems to perform various