A Survey on Evaluation of LLMs
Free while signed in. Answers cite the passages they came from.

A comprehensive overview of evaluation methods covering what, where, and how to evaluate LLMs.
Three-axis taxonomy: Organizes evaluation along what-to-evaluate (NLP tasks, robustness, ethics, trustworthiness), where-to-evaluate (benchmarks, datasets), and how-to-evaluate (automatic, human, LLM-as-judge).
Benchmark catalog: Surveys the major benchmarks of 2023 including MMLU, HELM, BIG-bench, and AgentBench with strengths and limitations.
Failure-mode analysis: Documents where current evaluations fall short - contamination, saturation, prompt sensitivity, and lack of task diversity.
Evaluation field primer: Became a standard citation for researchers entering LLM evaluation, helping formalize the sub-field.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack