LLM Alignment Survey
First page

Paper summary
A comprehensive survey of LLM alignment research spanning theoretical foundations to adversarial pressure.
Ask this paper
01
Outer and inner alignment: Distinguishes outer alignment (specifying the right objective) from inner alignment (ensuring the model actually pursues that objective).
02
Mechanistic interpretability: Reviews interpretability as an alignment tool, covering circuits, activation patching, and probing approaches.
03
Adversarial pressure: Catalogs known attacks on aligned LLMs including jailbreaks, prompt injection, and reward hacking.
04
Evaluation and directions: Discusses alignment evaluation methodologies and open problems, including scalable oversight for future systems beyond human capability.