🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training · Memory

YaRN (Efficient Context Extension)

Free while signed in. Answers cite the passages they came from.

First page
YaRN (Efficient Context Extension)
The curator’s take

YaRN is a compute-efficient method for extending the context window of LLMs well beyond their pretrained length.

Key points
01

Rotary-embedding scaling: Extends RoPE-based context length via a combined attention and NTK-aware scaling scheme, avoiding the degradation of naive interpolation.

02

Fine-tune extrapolation: Extrapolates meaningfully beyond the limited context seen during fine-tuning, so short fine-tune sequences can unlock much longer inference contexts.

03

128K context: Successfully scales Llama-family models to 128K-token context with minimal additional training compute.

04

Open recipe: Adopted widely across the open-source community as a standard recipe for extending Llama and other RoPE-based LLMs.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack