🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Evaluation · Reasoning

GAIA

First page
GAIA
Paper summary

Meta's GAIA is a benchmark for general AI assistants that requires reasoning, multimodal handling, web browsing, and tool use to solve real-world questions.

Ask this paper

Key points
01

Real-world questions: Questions are conceptually simple for humans but require integrated reasoning, web research, and tool use - a realistic test for assistant-style AI.

02

Massive human-model gap: Humans achieve 92% accuracy while GPT-4 with plugins achieves only 15% - the widest human-AI gap on any major 2023 benchmark.

03

Level-graduated difficulty: Three difficulty levels let researchers measure incremental progress rather than just binary success/failure.

04

Agent-first evaluation: Explicitly designed to test AI assistants, not base LLMs - a framing that has since become dominant for agent evaluations.

Every Monday
Get next week’s papers.
Subscribe on Substack