🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation

GDPval

First page
GDPval
Paper summary

GDPval is a new benchmark of 1,320 real-world tasks across 44 occupations in 9 major GDP sectors, graded by industry experts with a 220-task gold set. It shows frontier models improve roughly linearly and are nearing expert parity, with Claude Opus 4.1 preferred or tied 47.6% of the time, while GPT-5 leads in accuracy. Model-plus-human workflows can reduce time and cost, and adding reasoning effort and prompt scaffolding further raises scores, with an open gold set and automated grader available for researchers.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack