🚀NEW LABGetting Started with Claude AgentsStart lab
Training

Distilling Step-by-Step!

First page
Distilling Step-by-Step!
Paper summary

A mechanism to train smaller models that outperform larger LLMs using fewer examples.

Ask this paper

Key points
01

Rationale extraction: Extracts CoT rationales from a larger teacher LLM, using them to augment smaller student model training.

02

Smaller beats larger: Distilled student models outperform LLMs 500x+ larger in size on benchmark reasoning tasks.

03

Data efficiency: Requires dramatically less labeled training data than standard fine-tuning by leveraging LLM rationales as free supervision.

04

Distillation paradigm: Influential for the 2024 proliferation of reasoning-distilled small models like Orca 2, Phi-3, and later reasoning-specific SLMs.

Every Monday
Get next week’s papers.
Subscribe on Substack