🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Memory · Multimodal

Claude 3

First page
Claude 3
Paper summary

Anthropic releases the Claude 3 family (Haiku, Sonnet, Opus), with Opus leapfrogging GPT-4 on many standard benchmarks and bringing frontier multimodal capability plus a much larger context window.

Ask this paper

Key points
01

Three tiers: Haiku for speed/cost, Sonnet as the balanced default, and Opus as the flagship; each covers analysis, forecasting, content creation, code, and multilingual translation (Spanish, Japanese, French, etc.).

02

Benchmark leadership: Opus posts 86.8% on MMLU and 84.9% on HumanEval, edging past GPT-4 on reasoning and code benchmarks while also leading on MATH and GSM8K.

03

Long context: All three models ship with a 200K-token context window, extensible to 1M tokens for select customers, targeting long-document and long-agent-trajectory use cases.

04

Vision and refusals: Strong vision for photos, charts, and graphs; Anthropic reports more nuanced request handling with materially fewer unwarranted refusals compared to prior Claude generations.

Every Monday
Get next week’s papers.
Subscribe on Substack