🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training

Distilling Step-by-Step!

Free while signed in. Answers cite the passages they came from.

First page
Distilling Step-by-Step!
The curator’s take

A mechanism to train smaller models that outperform larger LLMs using fewer examples.

Key points
01

Rationale extraction: Extracts CoT rationales from a larger teacher LLM, using them to augment smaller student model training.

02

Smaller beats larger: Distilled student models outperform LLMs 500x+ larger in size on benchmark reasoning tasks.

03

Data efficiency: Requires dramatically less labeled training data than standard fine-tuning by leveraging LLM rationales as free supervision.

04

Distillation paradigm: Influential for the 2024 proliferation of reasoning-distilled small models like Orca 2, Phi-3, and later reasoning-specific SLMs.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack