🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

Extracting Interpretable Features from Claude 3 Sonnet

Paper preview
Extracting Interpretable Features from Claude 3 Sonnet
Paper summary

presents an effective method to extract millions of abstract features from an LLM that represent specific concepts; these concepts could represent people, places, programming abstractions, emotion, and more; reports that some of the discovered features are directly related to the safety aspects of the model; finds features directly related to security vulnerabilities and backdoors in code, bias, deception, sycophancy; and dangerous/criminal content, and more; these features are also used to intuititively steer the model’s output.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack