Open Problems and Limitations of RLHF
First page

Paper summary
A comprehensive survey of open problems and fundamental limitations of RLHF as an alignment approach.
Ask this paper
01
Scope: Catalogs issues across the entire RLHF pipeline - preference data collection, reward modeling, policy optimization, and evaluation.
02
Fundamental limitations: Discusses issues that can't be solved by incremental engineering alone, including the difficulty of specifying human preferences completely.
03
Reward hacking taxonomy: Organizes the many varieties of reward hacking seen in practice, from sycophancy to specification gaming.
04
Research agenda: Argues for investment in alignment approaches beyond RLHF that can address its structural limitations - a precursor to DPO and related methods.