🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning

Open Problems and Limitations of RLHF

First page
Open Problems and Limitations of RLHF
Paper summary

A comprehensive survey of open problems and fundamental limitations of RLHF as an alignment approach.

Ask this paper

Key points
01

Scope: Catalogs issues across the entire RLHF pipeline - preference data collection, reward modeling, policy optimization, and evaluation.

02

Fundamental limitations: Discusses issues that can't be solved by incremental engineering alone, including the difficulty of specifying human preferences completely.

03

Reward hacking taxonomy: Organizes the many varieties of reward hacking seen in practice, from sycophancy to specification gaming.

04

Research agenda: Argues for investment in alignment approaches beyond RLHF that can address its structural limitations - a precursor to DPO and related methods.

Every Monday
Get next week’s papers.
Subscribe on Substack