🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

First page
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Paper summary

A benchmark using real human standardized exams.

Ask this paper

Key points
01

Real human exams: Uses actual college entrance exams, law school admission tests, math competitions, and civil service exams - not synthetic benchmarks.

02

Multilingual coverage: Includes English and Chinese versions of exams, testing bilingual capability.

03

Human-comparable scoring: Makes it natural to compare foundation models to human performance percentiles on identical exams.

04

Real-world evaluation: Became an important benchmark for claims about "expert-level" or "human-comparable" foundation model performance.

Every Monday
Get next week’s papers.
Subscribe on Substack