🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

LLM Explains Neurons in LLMs

Paper preview
LLM Explains Neurons in LLMs
Paper summary

OpenAI's automated interpretability pipeline using GPT-4 to explain GPT-2 neurons.

Ask this paper

Key points
01

GPT-4 as interpreter: Uses GPT-4 to generate natural-language explanations of what individual GPT-2 neurons detect.

02

Automated scoring: Also uses GPT-4 to score how well an explanation predicts the neuron's actual activations on new text.

03

Scale of interpretability: Enables scaling interpretability research to all neurons in a model, previously impractical with human effort.

04

Automated interpretability era: Sparked the automated interpretability research program that continued in 2024 with SAE-based techniques and Golden Gate Claude demos.

Every Monday
Get next week’s papers.
Subscribe on Substack