GPQA
Free while signed in. Answers cite the passages they came from.

A graduate-level Google-proof QA benchmark designed to stress-test reasoning in systems that might exceed human expertise.
448 expert questions: Consists of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry.
Google-proof by design: Questions are constructed so that even with unrestricted internet access, non-experts (~34%) perform only slightly better than random on them.
GPT-4 gets 39%: The strongest GPT-4 baseline hits only 39% accuracy, showing a clear headroom for frontier models on expert-level reasoning.
Scalable oversight testbed: Explicitly designed to enable scalable oversight research - experiments in supervising models whose knowledge may exceed the supervisors'.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack