🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning · Evaluation

GPQA

First page
GPQA
Paper summary

A graduate-level Google-proof QA benchmark designed to stress-test reasoning in systems that might exceed human expertise.

Ask this paper

Key points
01

448 expert questions: Consists of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry.

02

Google-proof by design: Questions are constructed so that even with unrestricted internet access, non-experts (~34%) perform only slightly better than random on them.

03

GPT-4 gets 39%: The strongest GPT-4 baseline hits only 39% accuracy, showing a clear headroom for frontier models on expert-level reasoning.

04

Scalable oversight testbed: Explicitly designed to enable scalable oversight research - experiments in supervising models whose knowledge may exceed the supervisors'.

Every Monday
Get next week’s papers.
Subscribe on Substack