Phi-3
Free while signed in. Answers cite the passages they came from.

Microsoft's Phi-3 is a family of small language models (3.8B, 7B, 14B) trained on 3.3-4.8T tokens of heavily filtered web data combined with synthetic data. The flagship phi-3-mini rivals Mixtral 8x7B and GPT-3.5 while being small enough to run locally on a phone.
Data quality over scale: Training prioritizes curated web and synthetic data over raw token volume, suggesting that data quality is the main lever for small-model capability rather than sheer parameter count.
Benchmark results: phi-3-mini reaches 69% on MMLU and 8.38 on MT-bench; phi-3-small hits 75% on MMLU and phi-3-medium hits 78%, closing much of the gap with models an order of magnitude larger.
Long-context variant: A phi-3-mini-128K version extends the default 4K window to 128K tokens while preserving the base model's quality.
On-device deployment: The 3.8B mini model can be quantized and deployed on an iPhone 14, making it one of the first genuinely usable LLMs at that form factor.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack