🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Memory

LongLLaMA

Free while signed in. Answers cite the passages they came from.

First page
LongLLaMA
The curator’s take

Extends LLaMA's context length using a contrastive training process that reshapes the (key, value) space.

Key points
01

Focused Transformer: Uses contrastive training to make memory-augmented attention more discriminative, reducing distraction from irrelevant context.

02

Length extrapolation: Demonstrates long-context capability well beyond the original LLaMA 2K/4K window through its memory mechanism.

03

Long-context tasks: Shows improvements on passkey retrieval and long-form summarization tasks that stress long-range attention.

04

Efficient extension: Part of the 2023 explosion of context-window-extension techniques that would culminate in ~1M-token proprietary models the following year.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack