🚀NEW LABGetting Started with Claude AgentsStart lab
Safety · Reinforcement Learning

Survey of Aligned LLMs

First page
Survey of Aligned LLMs
Paper summary

A comprehensive overview of alignment approaches covering data, training, and evaluation.

Ask this paper

Key points
01

Full-stack view: Covers preference data collection, RLHF variants, DPO-style direct methods, and alignment evaluation in one unified reference.

02

Taxonomy of methods: Organizes alignment techniques into clear families (outer alignment vs. inner alignment, value alignment vs. behavior alignment).

03

Practical pitfalls: Documents known failure modes like reward hacking, sycophancy, and mode collapse that practitioners should watch for.

04

Reference document: Frequently cited in alignment onboarding material as the first-pass overview for new researchers.

Every Monday
Get next week’s papers.
Subscribe on Substack