🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation

Reliability without Validity

First page
Reliability without Validity
Paper summary

LLM-as-a-Judge is the default way to evaluate language models, but validating those judges with exact-match agreement never corrects for chance and systematically overstates how good they are. In the largest audit to date, spanning 21 judges from nine providers across MT-Bench, JudgeBench, and RewardBench over 118 runs and roughly 541,000 judgments, the gap between raw agreement and chance-corrected Cohen's kappa runs 33 to 41 percentage points, rankings shift by up to 14 positions across benchmarks, and high test-retest reliability coexists with severe position bias. The authors distill their findings into a Minimum Viable Validation Protocol so teams can stress-test judges before trusting them.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack