LLM Explains Neurons in LLMs
Free while signed in. Answers cite the passages they came from.

OpenAI's automated interpretability pipeline using GPT-4 to explain GPT-2 neurons.
GPT-4 as interpreter: Uses GPT-4 to generate natural-language explanations of what individual GPT-2 neurons detect.
Automated scoring: Also uses GPT-4 to score how well an explanation predicts the neuron's actual activations on new text.
Scale of interpretability: Enables scaling interpretability research to all neurons in a model, previously impractical with human effort.
Automated interpretability era: Sparked the automated interpretability research program that continued in 2024 with SAE-based techniques and Golden Gate Claude demos.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack