🚀NEW LABGetting Started with Claude AgentsStart lab
Data

Beyond Human Data (ReST-EM)

First page
Beyond Human Data (ReST-EM)
Paper summary

DeepMind's ReST-EM shows that model-generated data plus a reward function can substantially reduce dependence on human-generated data.

Ask this paper

Key points
01

Expectation-Maximization framing: Generates candidate solutions from the current model, filters using a reward/verifier, and fine-tunes on the filtered set - repeat.

02

Verifiable rewards: Uses automatic verifiers (e.g., correct-answer checks) as the reward signal, sidestepping the need for a learned reward model on scarce tasks.

03

PaLM 2 gains: Scales effectively on PaLM 2 for math and code tasks, outperforming standard SFT on human data at matched compute.

04

Synthetic-data signal: A strong empirical case that self-generated filtered data can replace much of the human data bottleneck for reasoning tasks - a theme that grew through 2024.

Every Monday
Get next week’s papers.
Subscribe on Substack