Evaluating LLMs Survey
Free while signed in. Answers cite the passages they came from.

A comprehensive survey of LLM evaluation covering benchmarks, methodologies, and open problems.
Task-wise organization: Organizes evaluation by task category - reasoning, knowledge, alignment, robustness, ethics, etc. - showing which benchmarks address which capabilities.
Automatic vs. human: Discusses the trade-offs between automatic metrics (cheap, inconsistent), LLM-as-a-Judge (scalable, biased), and human evaluation (reliable, expensive).
Contamination and robustness: Highlights contamination and robustness as cross-cutting concerns plaguing static benchmarks at all scales.
Frontier-model needs: Argues that evaluating frontier-scale LLMs requires new paradigms beyond simple benchmark accuracy, including interactive evaluation and behavioral testing.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack