How Language Models Use Long Contexts (Lost-in-the-Middle)
Free while signed in. Answers cite the passages they came from.

Shows LLM performance drops when relevant information is in the middle of a long context.
U-shaped performance curve: LMs perform best when relevant info is at the start or end of context, with substantial degradation for middle positions.
Cross-model phenomenon: Confirmed across GPT-3.5, GPT-4, Claude, and open-weight models - indicating a fundamental attention pattern rather than a bug.
QA and retrieval benchmarks: Demonstrated on multi-document QA and key-value retrieval tasks with varying context positions.
Foundational finding: Coined the phrase "lost in the middle" - one of the most widely-cited 2023 findings that shaped subsequent long-context benchmark and model design.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack