🚀NEW LABGetting Started with Claude AgentsStart lab
Memory · Training

Advancing Long-Context LLMs

First page
Advancing Long-Context LLMs
Paper summary

A survey of methodologies for improving Transformer long-context capability across pretraining, fine-tuning, and inference stages.

Ask this paper

Key points
01

Full-stack coverage: Organizes methods by training stage - pretraining objectives, position encoding, fine-tuning recipes, and inference-time interventions.

02

Position-encoding deep dive: Reviews RoPE variants, ALiBi, and other positional-encoding choices that dominate long-context extrapolation.

03

Efficient attention: Catalogs sparse, linear, and memory-augmented attention mechanisms that make longer contexts tractable.

04

Evaluation considerations: Addresses benchmark limitations including the "needle in a haystack" problem and the gap between nominal context length and effective usable context.

Every Monday
Get next week’s papers.
Subscribe on Substack