OpenAI's Automated Interpretability from paper "Language models can explain neurons in language models". Modified by Johnny Lin to add new models/context windows.
Default prompts from the main branch, strategy TokenActivationPair. Uses top 10 deduplicated activations.
Recent Explanations
comma tokens within software license/copyright headers, particularly ones appearing in standardized boilerplate sections about legal compliance or licensing terms.
technical terms and concepts in DynamoDB documentation or Q&A format, particularly activating on database terminology and structured question-answer formatting.
claude-4-5-sonnet
Dynamo DB<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n1. What is