🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Retrieval

Summary of a Haystack

First page
Summary of a Haystack
Paper summary

proposes a new task, SummHay, to test a model’s ability to process a Haystack and generate a summary that identifies the relevant insights and cites the source documents; reports that long-context LLMs score 20% on the benchmark which lags the human performance estimate (56%); RAG components is found to boost performance on the benchmark, which makes it a viable option for holistic RAG evaluation.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack