🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Safety

LLM Alignment Survey

Free while signed in. Answers cite the passages they came from.

First page
LLM Alignment Survey
The curator’s take

A comprehensive survey of LLM alignment research spanning theoretical foundations to adversarial pressure.

Key points
01

Outer and inner alignment: Distinguishes outer alignment (specifying the right objective) from inner alignment (ensuring the model actually pursues that objective).

02

Mechanistic interpretability: Reviews interpretability as an alignment tool, covering circuits, activation patching, and probing approaches.

03

Adversarial pressure: Catalogs known attacks on aligned LLMs including jailbreaks, prompt injection, and reward hacking.

04

Evaluation and directions: Discusses alignment evaluation methodologies and open problems, including scalable oversight for future systems beyond human capability.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack