YaRN (Efficient Context Extension)
First page

Paper summary
YaRN is a compute-efficient method for extending the context window of LLMs well beyond their pretrained length.
Ask this paper
01
Rotary-embedding scaling: Extends RoPE-based context length via a combined attention and NTK-aware scaling scheme, avoiding the degradation of naive interpolation.
02
Fine-tune extrapolation: Extrapolates meaningfully beyond the limited context seen during fine-tuning, so short fine-tune sequences can unlock much longer inference contexts.
03
128K context: Successfully scales Llama-family models to 128K-token context with minimal additional training compute.
04
Open recipe: Adopted widely across the open-source community as a standard recipe for extending Llama and other RoPE-based LLMs.