Secrets of RLHF in LLMs
First page

Paper summary
A deep investigation into RLHF with a focus on the inner workings of PPO, including open-source code.
Ask this paper
01
PPO internals exposed: Documents critical implementation details (reward normalization, advantage estimation, KL penalty scaling) that aren't in the original papers but make or break training.
02
Empirical ablations: Systematically studies which PPO components matter most, providing practical guidance for RLHF practitioners.
03
Open-source code: Releases a clean reference implementation that others can use to reproduce and iterate on RLHF.
04
RLHF demystification: Part of a broader 2023 wave demystifying RLHF, preparing the ground for simpler alternatives like DPO that arrived later that year.