Let's Verify Step by Step
First page

Paper summary
OpenAI's landmark paper on process reward models for mathematical reasoning.
Ask this paper
01
Process supervision: Rewards each correct step of reasoning rather than just the final answer, capturing partial credit and providing much denser training signal.
02
78% MATH solve rate: Achieves state-of-the-art on a representative subset of the MATH benchmark - a significant jump over outcome-reward baselines.
03
PRM800K dataset: Releases a massive dataset of 800K step-level correctness labels, enabling follow-up research on process reward models.
04
Reasoning revolution foundation: Directly influenced OpenAI's o1/o3 reasoning models and the broader 2024-25 push toward process-supervised reasoning.