Secrets of RLHF in LLMs
Free while signed in. Answers cite the passages they came from.

A deep investigation into RLHF with a focus on the inner workings of PPO, including open-source code.
PPO internals exposed: Documents critical implementation details (reward normalization, advantage estimation, KL penalty scaling) that aren't in the original papers but make or break training.
Empirical ablations: Systematically studies which PPO components matter most, providing practical guidance for RLHF practitioners.
Open-source code: Releases a clean reference implementation that others can use to reproduce and iterate on RLHF.
RLHF demystification: Part of a broader 2023 wave demystifying RLHF, preparing the ground for simpler alternatives like DPO that arrived later that year.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack