Direct Preference Optimization (DPO)
First page

Paper summary
Rafailov et al.'s simpler alternative to RLHF that rivals full RL-based alignment.
Ask this paper
01
Classification, not RL: Reformulates preference learning as a classification problem on preference pairs, skipping the complex RL loop entirely.
02
Theoretical equivalence: Mathematically equivalent to RLHF under certain assumptions, extracting the implicit reward function directly.
03
Training stability: Much more stable and hyperparameter-robust than PPO-based RLHF, dramatically lowering the barrier to entry.
04
Industry-wide adoption: Became the default alignment method throughout 2024 (Zephyr, Tulu, Llama 3 pipelines) and ushered in the era of RL-free preference optimization.