🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

LLM Alignment Survey

First page
LLM Alignment Survey
Paper summary

A comprehensive survey of LLM alignment research spanning theoretical foundations to adversarial pressure.

Ask this paper

Key points
01

Outer and inner alignment: Distinguishes outer alignment (specifying the right objective) from inner alignment (ensuring the model actually pursues that objective).

02

Mechanistic interpretability: Reviews interpretability as an alignment tool, covering circuits, activation patching, and probing approaches.

03

Adversarial pressure: Catalogs known attacks on aligned LLMs including jailbreaks, prompt injection, and reward hacking.

04

Evaluation and directions: Discusses alignment evaluation methodologies and open problems, including scalable oversight for future systems beyond human capability.

Every Monday
Get next week’s papers.
Subscribe on Substack