Phi-3

Microsoft's Phi-3 is a family of small language models (3.8B, 7B, 14B) trained on 3.3-4.8T tokens of heavily filtered web data combined with synthetic data. The flagship phi-3-mini rivals Mixtral 8x7B and GPT-3.5 while being small enough to run locally on a phone.
Ask this paper
Data quality over scale: Training prioritizes curated web and synthetic data over raw token volume, suggesting that data quality is the main lever for small-model capability rather than sheer parameter count.
Benchmark results: phi-3-mini reaches 69% on MMLU and 8.38 on MT-bench; phi-3-small hits 75% on MMLU and phi-3-medium hits 78%, closing much of the gap with models an order of magnitude larger.
Long-context variant: A phi-3-mini-128K version extends the default 4K window to 128K tokens while preserving the base model's quality.
On-device deployment: The 3.8B mini model can be quantized and deployed on an iPhone 14, making it one of the first genuinely usable LLMs at that form factor.