AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
First page

Paper summary
A benchmark using real human standardized exams.
Ask this paper
01
Real human exams: Uses actual college entrance exams, law school admission tests, math competitions, and civil service exams - not synthetic benchmarks.
02
Multilingual coverage: Includes English and Chinese versions of exams, testing bilingual capability.
03
Human-comparable scoring: Makes it natural to compare foundation models to human performance percentiles on identical exams.
04
Real-world evaluation: Became an important benchmark for claims about "expert-level" or "human-comparable" foundation model performance.