Recurrent Memory Finds What LLMs Miss
Free while signed in. Answers cite the passages they came from.

Introduces BABILong, a new long-context benchmark, and shows that transformers with recurrent memory can handle sequences far beyond vanilla LLMs.
BABILong benchmark: Extends the bAbI reasoning tasks by embedding them inside arbitrarily long distractor text, stress-testing whether models actually use their context or ignore most of it.
Attention heavy tails: On BABILong, GPT-4 and RAG systems effectively rely on the first ~25% of the input - a stark illustration that "100K context" does not mean "100K useful context".
Recurrent memory wins: Augmenting a GPT-2-scale transformer with recurrent memory lets it process sequences up to ~11M tokens while still answering the embedded questions correctly.
Direction signal: Argues recurrent memory is a simpler and cheaper path to genuinely long-context models than ever-larger attention windows.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack