🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training

o1 Replication Journey - Part 2

Free while signed in. Answers cite the passages they came from.

First page
o1 Replication Journey - Part 2
The curator’s take

shows that combining simple distillation from o1's API with supervised fine-tuning significantly boosts performance on complex math reasoning tasks; a base model fine-tuned on simply tens of thousands of samples o1-distilled long-thought chains outperform o1-preview on the American Invitational Mathematics Examination (AIME).

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack