Overview of LLMs for Evaluation
Free while signed in. Answers cite the passages they came from.

A thorough survey of LLM-as-a-Judge and LLM-based evaluation methodologies, mapping strengths, limitations, and open problems.
Taxonomy: Groups evaluators by whether they use prompt engineering alone, calibrated prompts, or fine-tuned open-source LLMs - a clean mental model for practitioners choosing a judge.
Task coverage: Analyzes LLM evaluators across summarization, dialogue, translation, code, and reasoning tasks, showing where they are and aren't reliable.
Failure modes: Reviews known biases including length bias, position bias, self-preference, and sycophancy - and the mitigation strategies published for each.
Future directions: Argues for standardized benchmarks for evaluators themselves (meta-evaluation) and more transparent calibration procedures.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack