🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Evaluation · Memory · Multimodal

Claude 3

Free while signed in. Answers cite the passages they came from.

First page
Claude 3
The curator’s take

Anthropic releases the Claude 3 family (Haiku, Sonnet, Opus), with Opus leapfrogging GPT-4 on many standard benchmarks and bringing frontier multimodal capability plus a much larger context window.

Key points
01

Three tiers: Haiku for speed/cost, Sonnet as the balanced default, and Opus as the flagship; each covers analysis, forecasting, content creation, code, and multilingual translation (Spanish, Japanese, French, etc.).

02

Benchmark leadership: Opus posts 86.8% on MMLU and 84.9% on HumanEval, edging past GPT-4 on reasoning and code benchmarks while also leading on MATH and GSM8K.

03

Long context: All three models ship with a 200K-token context window, extensible to 1M tokens for select customers, targeting long-document and long-agent-trajectory use cases.

04

Vision and refusals: Strong vision for photos, charts, and graphs; Anthropic reports more nuanced request handling with materially fewer unwarranted refusals compared to prior Claude generations.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack