🚀NEW LABGetting Started with Claude AgentsStart lab
Training

o1 Replication Journey - Part 2

First page
o1 Replication Journey - Part 2
Paper summary

shows that combining simple distillation from o1's API with supervised fine-tuning significantly boosts performance on complex math reasoning tasks; a base model fine-tuned on simply tens of thousands of samples o1-distilled long-thought chains outperform o1-preview on the American Invitational Mathematics Examination (AIME).

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack