LLM Explains Neurons in LLMs
Paper preview

Paper summary
OpenAI's automated interpretability pipeline using GPT-4 to explain GPT-2 neurons.
Ask this paper
01
GPT-4 as interpreter: Uses GPT-4 to generate natural-language explanations of what individual GPT-2 neurons detect.
02
Automated scoring: Also uses GPT-4 to score how well an explanation predicts the neuron's actual activations on new text.
03
Scale of interpretability: Enables scaling interpretability research to all neurons in a model, previously impractical with human effort.
04
Automated interpretability era: Sparked the automated interpretability research program that continued in 2024 with SAE-based techniques and Golden Gate Claude demos.