🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Reasoning

Imitating Reasoning Process of Larger LLMs (Orca)

Free while signed in. Answers cite the passages they came from.

First page
Imitating Reasoning Process of Larger LLMs (Orca)
The curator’s take

Microsoft's 13B model that imitates GPT-4's reasoning traces.

Key points
01

Explanation tuning: Trains on detailed step-by-step explanations from GPT-4, not just final answers - capturing the reasoning process.

02

Scale and diversity: Leverages millions of diverse imitation examples spanning reasoning tasks, dialogue, and instruction-following.

03

Beats Vicuna-13B: Surpasses instruction-tuned Vicuna-13B in zero-shot reasoning, demonstrating explanation-data quality matters.

04

Small-model reasoning: Kicked off a line of research on reasoning distillation that would continue through Orca 2 and into 2024's reasoning-specific SLMs.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack