Advancing Long-Context LLMs
Free while signed in. Answers cite the passages they came from.

A survey of methodologies for improving Transformer long-context capability across pretraining, fine-tuning, and inference stages.
Full-stack coverage: Organizes methods by training stage - pretraining objectives, position encoding, fine-tuning recipes, and inference-time interventions.
Position-encoding deep dive: Reviews RoPE variants, ALiBi, and other positional-encoding choices that dominate long-context extrapolation.
Efficient attention: Catalogs sparse, linear, and memory-augmented attention mechanisms that make longer contexts tractable.
Evaluation considerations: Addresses benchmark limitations including the "needle in a haystack" problem and the gap between nominal context length and effective usable context.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack