🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Evaluation · Agents

Agents' Last Exam

Free while signed in. Answers cite the passages they came from.

First page
Agents' Last Exam
The curator’s take

From Berkeley RDI, Agents' Last Exam (ALE) is a living benchmark built to measure whether agents can do economically valuable work, not just score well on academic tests. It was assembled with more than 250 industry experts and maps over 1,000 verifiable tasks to the U.S. federal occupational taxonomy, organized as 55 subfields across 13 industry clusters. Every task has an objective, checkable outcome, so there is no subjective human grading, and the pool is designed to keep growing as new workflows are onboarded.

Key points
01

Grounded in real occupations: Tasks are defined against O*NET and SOC 2018 and span non-physical industries, deliberately targeting the professional workflows where agents would actually be deployed rather than puzzle-style problems.

02

Three difficulty tiers: Work is split into Near-Term, Full-Spectrum, and Last-Exam tiers, letting the benchmark track both near-term usefulness and the long tail of hard, multi-step jobs.

03

Far from saturated: The hardest tier sits at just a 2.6% average full pass rate across mainstream harnesses, and even strong setups like Codex with GPT-5.5 score below 50% on the easiest tier and under 10% on the hardest.

04

Why it matters: Strong scores on existing benchmarks have not translated into economically meaningful deployment. ALE reframes evaluation around verifiable, expert-curated work, giving a moving target that should resist saturation as agents improve.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack