NaturalThoughts
Free while signed in. Answers cite the passages they came from.

This paper introduces NaturalThoughts, a large-scale dataset of reasoning traces distilled from DeepSeek-R1 using questions from the NaturalReasoning corpus. It challenges the "Less is More" hypothesis by showing that simply scaling up high-quality reasoning traces, without aggressive filtering, yields robust and general improvements across STEM reasoning tasks for smaller models like Llama-3.1-8B and Qwen-2.5-7B.
Scale beats sparsity: Training on hundreds of thousands of randomly selected reasoning traces from NaturalThoughts consistently outperforms curated datasets like LIMO and S1K on GPQA-Diamond, MMLU-Pro, and SuperGPQA, reversing the prior trend where small, curated datasets showed outsized benefits.
Hard examples help more: Selecting training data by difficulty, e.g., examples with model disagreement or long reasoning chains, leads to greater sample efficiency than random sampling. The most effective subsets use disagreement between teacher models as a proxy for reasoning complexity.
Diversity matters, but not in an obvious way: Contrary to expectations, semantic or topical diversity in question domains yielded less gain than diversity in reasoning strategies themselves. Traces with a mix of tactics like self-verification, backtracking, and synthesis provided stronger generalization across tasks.
Mixed distillation improves efficiency: Blending full CoT traces (System-2) with final answers only (System-1) enables inference-time control over reasoning length. Difficulty-based mixing, applying System-2 only for harder examples, beats both random mixing and pure System-2, achieving higher accuracy with fewer tokens at test time.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack