AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsA benchmark using real human standardized exams.
01
Real human exams: Uses actual college entrance exams, law school admission tests, math competitions, and civil service exams - not synthetic benchmarks.
02
Multilingual coverage: Includes English and Chinese versions of exams, testing bilingual capability.
03
Human-comparable scoring: Makes it natural to compare foundation models to human performance percentiles on identical exams.
04
Real-world evaluation: Became an important benchmark for claims about "expert-level" or "human-comparable" foundation model performance.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack