A Survey on Evaluation of LLMs
First page

Paper summary
A comprehensive overview of evaluation methods covering what, where, and how to evaluate LLMs.
Ask this paper
01
Three-axis taxonomy: Organizes evaluation along what-to-evaluate (NLP tasks, robustness, ethics, trustworthiness), where-to-evaluate (benchmarks, datasets), and how-to-evaluate (automatic, human, LLM-as-judge).
02
Benchmark catalog: Surveys the major benchmarks of 2023 including MMLU, HELM, BIG-bench, and AgentBench with strengths and limitations.
03
Failure-mode analysis: Documents where current evaluations fall short - contamination, saturation, prompt sensitivity, and lack of task diversity.
04
Evaluation field primer: Became a standard citation for researchers entering LLM evaluation, helping formalize the sub-field.