🚀NEW LABGetting Started with Claude AgentsStart lab
Data · Efficiency

Phi-3

First page
Phi-3
Paper summary

Microsoft's Phi-3 is a family of small language models (3.8B, 7B, 14B) trained on 3.3-4.8T tokens of heavily filtered web data combined with synthetic data. The flagship phi-3-mini rivals Mixtral 8x7B and GPT-3.5 while being small enough to run locally on a phone.

Ask this paper

Key points
01

Data quality over scale: Training prioritizes curated web and synthetic data over raw token volume, suggesting that data quality is the main lever for small-model capability rather than sheer parameter count.

02

Benchmark results: phi-3-mini reaches 69% on MMLU and 8.38 on MT-bench; phi-3-small hits 75% on MMLU and phi-3-medium hits 78%, closing much of the gap with models an order of magnitude larger.

03

Long-context variant: A phi-3-mini-128K version extends the default 4K window to 128K tokens while preserving the base model's quality.

04

On-device deployment: The 3.8B mini model can be quantized and deployed on an iPhone 14, making it one of the first genuinely usable LLMs at that form factor.

Every Monday
Get next week’s papers.
Subscribe on Substack