🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Evaluation

A Survey on Evaluation of LLMs

Free while signed in. Answers cite the passages they came from.

First page
A Survey on Evaluation of LLMs
The curator’s take

A comprehensive overview of evaluation methods covering what, where, and how to evaluate LLMs.

Key points
01

Three-axis taxonomy: Organizes evaluation along what-to-evaluate (NLP tasks, robustness, ethics, trustworthiness), where-to-evaluate (benchmarks, datasets), and how-to-evaluate (automatic, human, LLM-as-judge).

02

Benchmark catalog: Surveys the major benchmarks of 2023 including MMLU, HELM, BIG-bench, and AgentBench with strengths and limitations.

03

Failure-mode analysis: Documents where current evaluations fall short - contamination, saturation, prompt sensitivity, and lack of task diversity.

04

Evaluation field primer: Became a standard citation for researchers entering LLM evaluation, helping formalize the sub-field.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack