L-Eval
First page

Paper summary
A standardized evaluation suite for long-context language models.
Ask this paper
01
Dataset scale: 411 long documents covering over 2K query-response pairs across law, finance, school lectures, long conversations, novels, and meetings.
02
Realistic domains: Moves beyond synthetic needle-in-haystack tests toward practical long-form applications users actually encounter.
03
Evaluation methodology: Provides multiple evaluation protocols including exact match, n-gram, and LLM-as-judge to cross-validate results.
04
Long-context benchmark: Became a reference benchmark during 2023's context-window race, paving the way for later benchmarks like LongBench and RULER.