Textbooks Are All You Need (phi-1)
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsIntroduces a 1.3B parameter code LLM trained on textbook-quality data.
01
Data-quality thesis: Trained on a curated selection of textbook-quality web data plus synthetic textbooks/exercises generated with GPT-3.5.
02
Small model, strong HumanEval: Achieves 50.6% pass@1 on HumanEval despite being 1.3B - beating much larger models on code generation.
03
4-day training: Trained in just 4 days on 8 A100s, showing that aggressive data selection can substitute for massive compute.
04
Phi-series launch: Kicked off Microsoft's Phi-series (Phi-1.5, Phi-2, Phi-3) and catalyzed the "small-but-smart" model research program.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack