🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation

Easy-to-Hard Generalization

First page
Easy-to-Hard Generalization
Paper summary

UNC researchers show that LLMs often generalize well from easy training data to hard evaluation data, with implications for scalable oversight.

Ask this paper

Key points
01

Surprising finding: Fine-tuning on *easy* examples can match or beat fine-tuning on *hard* examples when the evaluation is on hard examples - counter to the usual "train on the distribution you care about" intuition.

02

Broad empirical support: Shown across a range of tasks including math, reasoning, and knowledge-intensive QA, using multiple model families and fine-tuning regimes.

03

Implication for scalable oversight: If humans only supervise easy examples and the model still generalizes to hard ones, scalable oversight for superhuman models may be more tractable than previously assumed.

04

Caveats: The effect depends on task structure and easy/hard defined by human-calibrated difficulty - not all gradations of difficulty admit this generalization.

Every Monday
Get next week’s papers.
Subscribe on Substack