L-Eval
Free while signed in. Answers cite the passages they came from.

A standardized evaluation suite for long-context language models.
Dataset scale: 411 long documents covering over 2K query-response pairs across law, finance, school lectures, long conversations, novels, and meetings.
Realistic domains: Moves beyond synthetic needle-in-haystack tests toward practical long-form applications users actually encounter.
Evaluation methodology: Provides multiple evaluation protocols including exact match, n-gram, and LLM-as-judge to cross-validate results.
Long-context benchmark: Became a reference benchmark during 2023's context-window race, paving the way for later benchmarks like LongBench and RULER.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack