GPQA
First page

Paper summary
A graduate-level Google-proof QA benchmark designed to stress-test reasoning in systems that might exceed human expertise.
Ask this paper
01
448 expert questions: Consists of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry.
02
Google-proof by design: Questions are constructed so that even with unrestricted internet access, non-experts (~34%) perform only slightly better than random on them.
03
GPT-4 gets 39%: The strongest GPT-4 baseline hits only 39% accuracy, showing a clear headroom for frontier models on expert-level reasoning.
04
Scalable oversight testbed: Explicitly designed to enable scalable oversight research - experiments in supervising models whose knowledge may exceed the supervisors'.