Distilling Step-by-Step!
Free while signed in. Answers cite the passages they came from.

A mechanism to train smaller models that outperform larger LLMs using fewer examples.
Rationale extraction: Extracts CoT rationales from a larger teacher LLM, using them to augment smaller student model training.
Smaller beats larger: Distilled student models outperform LLMs 500x+ larger in size on benchmark reasoning tasks.
Data efficiency: Requires dramatically less labeled training data than standard fine-tuning by leveraging LLM rationales as free supervision.
Distillation paradigm: Influential for the 2024 proliferation of reasoning-distilled small models like Orca 2, Phi-3, and later reasoning-specific SLMs.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack