Extending Context Window of LLMs (PI)
First page

Paper summary
Position Interpolation extends LLaMA's context to 32K with minimal fine-tuning (within 1000 steps).
Ask this paper
01
Position interpolation: Linearly interpolates positional indices so pretrained RoPE attention generalizes to longer sequences without breaking.
02
1000-step adaptation: Requires only ~1000 fine-tuning steps versus prior methods that needed much more compute.
03
Quality preservation: Maintains strong performance on tasks while reaching 32K context - both long-context tasks and standard-length benchmarks.
04
Standard long-context recipe: Became the standard approach for extending open-source model context windows throughout 2023 and early 2024.