🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning

Imitating Reasoning Process of Larger LLMs (Orca)

First page
Imitating Reasoning Process of Larger LLMs (Orca)
Paper summary

Microsoft's 13B model that imitates GPT-4's reasoning traces.

Ask this paper

Key points
01

Explanation tuning: Trains on detailed step-by-step explanations from GPT-4, not just final answers - capturing the reasoning process.

02

Scale and diversity: Leverages millions of diverse imitation examples spanning reasoning tasks, dialogue, and instruction-following.

03

Beats Vicuna-13B: Surpasses instruction-tuned Vicuna-13B in zero-shot reasoning, demonstrating explanation-data quality matters.

04

Small-model reasoning: Kicked off a line of research on reasoning distillation that would continue through Orca 2 and into 2024's reasoning-specific SLMs.

Every Monday
Get next week’s papers.
Subscribe on Substack