LegalBench
Free while signed in. Answers cite the passages they came from.

A collaboratively constructed benchmark for measuring legal reasoning in LLMs.
162 tasks: Covers 162 legal-reasoning tasks designed by legal experts, significantly broader than prior legal benchmarks.
Six reasoning categories: Categorizes tasks across rule-recall, rule-application, rule-conclusion, interpretation, rhetorical-analysis, and issue-spotting.
Collaborative construction: Built through collaboration with legal practitioners to ensure tasks reflect real legal reasoning rather than generic NLP tasks dressed in legal vocabulary.
LLM-lawyer evaluation: Provides the first rigorous benchmark for systematically evaluating LLM legal capability - essential for responsible deployment in legal workflows.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack