🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Memory

Long-context LLMs Struggle with Long In-Context Learning

Free while signed in. Answers cite the passages they came from.

First page
Long-context LLMs Struggle with Long In-Context Learning
The curator’s take

LongICLBench stress-tests 13 long-context LLMs on extreme-label classification with up to 174 classes and 50K-token prompts, exposing sharp quality cliffs beyond 20K tokens.

Key points
01

Benchmark design: Tasks cover 28-174 label classes packed into 2K-50K token contexts, requiring models to integrate many demonstrations rather than retrieve a single fact.

02

Short-to-mid context OK: Most models perform adequately under ~20K tokens, suggesting that advertised long-context windows can at least hold instructions and examples.

03

Sharp degradation past 20K: On the hardest task (Discovery, 174 labels) every open model collapses, and all but GPT-4 dip dramatically as context grows beyond 20K tokens.

04

Position bias: Models over-predict labels that appear later in the prompt, revealing that long-context ICL fails both on reasoning across examples and on treating positions symmetrically.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack