🚀NEW LABGetting Started with Claude AgentsStart lab
Memory

LongLLaMA

First page
LongLLaMA
Paper summary

Extends LLaMA's context length using a contrastive training process that reshapes the (key, value) space.

Ask this paper

Key points
01

Focused Transformer: Uses contrastive training to make memory-augmented attention more discriminative, reducing distraction from irrelevant context.

02

Length extrapolation: Demonstrates long-context capability well beyond the original LLaMA 2K/4K window through its memory mechanism.

03

Long-context tasks: Shows improvements on passkey retrieval and long-form summarization tasks that stress long-range attention.

04

Efficient extension: Part of the 2023 explosion of context-window-extension techniques that would culminate in ~1M-token proprietary models the following year.

Every Monday
Get next week’s papers.
Subscribe on Substack