Claude 3
Free while signed in. Answers cite the passages they came from.

Anthropic releases the Claude 3 family (Haiku, Sonnet, Opus), with Opus leapfrogging GPT-4 on many standard benchmarks and bringing frontier multimodal capability plus a much larger context window.
Three tiers: Haiku for speed/cost, Sonnet as the balanced default, and Opus as the flagship; each covers analysis, forecasting, content creation, code, and multilingual translation (Spanish, Japanese, French, etc.).
Benchmark leadership: Opus posts 86.8% on MMLU and 84.9% on HumanEval, edging past GPT-4 on reasoning and code benchmarks while also leading on MATH and GSM8K.
Long context: All three models ship with a 200K-token context window, extensible to 1M tokens for select customers, targeting long-document and long-agent-trajectory use cases.
Vision and refusals: Strong vision for photos, charts, and graphs; Anthropic reports more nuanced request handling with materially fewer unwarranted refusals compared to prior Claude generations.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack