Survey of Aligned LLMs
Free while signed in. Answers cite the passages they came from.

A comprehensive overview of alignment approaches covering data, training, and evaluation.
Full-stack view: Covers preference data collection, RLHF variants, DPO-style direct methods, and alignment evaluation in one unified reference.
Taxonomy of methods: Organizes alignment techniques into clear families (outer alignment vs. inner alignment, value alignment vs. behavior alignment).
Practical pitfalls: Documents known failure modes like reward hacking, sycophancy, and mode collapse that practitioners should watch for.
Reference document: Frequently cited in alignment onboarding material as the first-pass overview for new researchers.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack