Claude 3

Anthropic releases the Claude 3 family (Haiku, Sonnet, Opus), with Opus leapfrogging GPT-4 on many standard benchmarks and bringing frontier multimodal capability plus a much larger context window.
Ask this paper
Three tiers: Haiku for speed/cost, Sonnet as the balanced default, and Opus as the flagship; each covers analysis, forecasting, content creation, code, and multilingual translation (Spanish, Japanese, French, etc.).
Benchmark leadership: Opus posts 86.8% on MMLU and 84.9% on HumanEval, edging past GPT-4 on reasoning and code benchmarks while also leading on MATH and GSM8K.
Long context: All three models ship with a 200K-token context window, extensible to 1M tokens for select customers, targeting long-document and long-agent-trajectory use cases.
Vision and refusals: Strong vision for photos, charts, and graphs; Anthropic reports more nuanced request handling with materially fewer unwarranted refusals compared to prior Claude generations.