Direct Preference Optimization (DPO)
Free while signed in. Answers cite the passages they came from.

Rafailov et al.'s simpler alternative to RLHF that rivals full RL-based alignment.
Classification, not RL: Reformulates preference learning as a classification problem on preference pairs, skipping the complex RL loop entirely.
Theoretical equivalence: Mathematically equivalent to RLHF under certain assumptions, extracting the implicit reward function directly.
Training stability: Much more stable and hyperparameter-robust than PPO-based RLHF, dramatically lowering the barrier to entry.
Industry-wide adoption: Became the default alignment method throughout 2024 (Zephyr, Tulu, Llama 3 pipelines) and ushered in the era of RL-free preference optimization.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack