🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Reasoning · Evaluation

GPQA

Free while signed in. Answers cite the passages they came from.

First page
GPQA
The curator’s take

A graduate-level Google-proof QA benchmark designed to stress-test reasoning in systems that might exceed human expertise.

Key points
01

448 expert questions: Consists of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry.

02

Google-proof by design: Questions are constructed so that even with unrestricted internet access, non-experts (~34%) perform only slightly better than random on them.

03

GPT-4 gets 39%: The strongest GPT-4 baseline hits only 39% accuracy, showing a clear headroom for frontier models on expert-level reasoning.

04

Scalable oversight testbed: Explicitly designed to enable scalable oversight research - experiments in supervising models whose knowledge may exceed the supervisors'.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack