🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Evaluation

Easy-to-Hard Generalization

Free while signed in. Answers cite the passages they came from.

First page
Easy-to-Hard Generalization
The curator’s take

UNC researchers show that LLMs often generalize well from easy training data to hard evaluation data, with implications for scalable oversight.

Key points
01

Surprising finding: Fine-tuning on *easy* examples can match or beat fine-tuning on *hard* examples when the evaluation is on hard examples - counter to the usual "train on the distribution you care about" intuition.

02

Broad empirical support: Shown across a range of tasks including math, reasoning, and knowledge-intensive QA, using multiple model families and fine-tuning regimes.

03

Implication for scalable oversight: If humans only supervise easy examples and the model still generalizes to hard ones, scalable oversight for superhuman models may be more tractable than previously assumed.

04

Caveats: The effect depends on task structure and easy/hard defined by human-calibrated difficulty - not all gradations of difficulty admit this generalization.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack