Advancing Long-Context LLMs
First page

Paper summary
A survey of methodologies for improving Transformer long-context capability across pretraining, fine-tuning, and inference stages.
Ask this paper
01
Full-stack coverage: Organizes methods by training stage - pretraining objectives, position encoding, fine-tuning recipes, and inference-time interventions.
02
Position-encoding deep dive: Reviews RoPE variants, ALiBi, and other positional-encoding choices that dominate long-context extrapolation.
03
Efficient attention: Catalogs sparse, linear, and memory-augmented attention mechanisms that make longer contexts tractable.
04
Evaluation considerations: Addresses benchmark limitations including the "needle in a haystack" problem and the gap between nominal context length and effective usable context.