Llama 3
Free while signed in. Answers cite the passages they came from.

Meta's Llama 3 launches with 8B and 70B pretrained and instruction-tuned variants. Llama 3 8B beats Gemma 7B and Mistral 7B Instruct, and Llama 3 70B is competitive with Gemini Pro 1.5 and Claude 3 Sonnet on standard benchmarks.
Sizes and release: Meta ships 8B and 70B base and Instruct variants first; larger 400B+ models are still training and planned for later releases with multimodality and longer context.
Training data: Pretrained on 15T+ tokens (7x more than Llama 2) including 4x more code, with 5%+ non-English across 30+ languages. A new 128K-token tokenizer improves encoding efficiency by roughly 15%.
Benchmark results: The 70B Instruct model wins human preference rankings against Claude Sonnet, Mistral Medium, and GPT-3.5 across 12 use-case categories; the 8B sets a new state of the art for open models in its size class.
Open deployment: Weights are available across AWS, HuggingFace, Databricks, Google Cloud, Azure, and more, alongside a safety suite - Llama Guard 2, Code Shield, and CyberSec Eval 2 - for responsible deployment.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack