Easy-to-Hard Generalization

UNC researchers show that LLMs often generalize well from easy training data to hard evaluation data, with implications for scalable oversight.
Ask this paper
Surprising finding: Fine-tuning on *easy* examples can match or beat fine-tuning on *hard* examples when the evaluation is on hard examples - counter to the usual "train on the distribution you care about" intuition.
Broad empirical support: Shown across a range of tasks including math, reasoning, and knowledge-intensive QA, using multiple model families and fine-tuning regimes.
Implication for scalable oversight: If humans only supervise easy examples and the model still generalizes to hard ones, scalable oversight for superhuman models may be more tractable than previously assumed.
Caveats: The effect depends on task structure and easy/hard defined by human-calibrated difficulty - not all gradations of difficulty admit this generalization.