YaRN (Efficient Context Extension)
Free while signed in. Answers cite the passages they came from.

YaRN is a compute-efficient method for extending the context window of LLMs well beyond their pretrained length.
Rotary-embedding scaling: Extends RoPE-based context length via a combined attention and NTK-aware scaling scheme, avoiding the degradation of naive interpolation.
Fine-tune extrapolation: Extrapolates meaningfully beyond the limited context seen during fine-tuning, so short fine-tune sequences can unlock much longer inference contexts.
128K context: Successfully scales Llama-family models to 128K-token context with minimal additional training compute.
Open recipe: Adopted widely across the open-source community as a standard recipe for extending Llama and other RoPE-based LLMs.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack