🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation

Evaluating LLMs Survey

First page
Evaluating LLMs Survey
Paper summary

A comprehensive survey of LLM evaluation covering benchmarks, methodologies, and open problems.

Ask this paper

Key points
01

Task-wise organization: Organizes evaluation by task category - reasoning, knowledge, alignment, robustness, ethics, etc. - showing which benchmarks address which capabilities.

02

Automatic vs. human: Discusses the trade-offs between automatic metrics (cheap, inconsistent), LLM-as-a-Judge (scalable, biased), and human evaluation (reliable, expensive).

03

Contamination and robustness: Highlights contamination and robustness as cross-cutting concerns plaguing static benchmarks at all scales.

04

Frontier-model needs: Argues that evaluating frontier-scale LLMs requires new paradigms beyond simple benchmark accuracy, including interactive evaluation and behavioral testing.

Every Monday
Get next week’s papers.
Subscribe on Substack