🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

Extracting Concepts from GPT-4

Paper preview
Extracting Concepts from GPT-4
Paper summary

proposes a new scalable method based on sparse autoencoders to extract around 16 million interpretable patterns from GPT-4; the method demonstrates predictable scaling and is more efficient than previous techniques.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack