LongLLaMA
First page

Paper summary
Extends LLaMA's context length using a contrastive training process that reshapes the (key, value) space.
Ask this paper
01
Focused Transformer: Uses contrastive training to make memory-augmented attention more discriminative, reducing distraction from irrelevant context.
02
Length extrapolation: Demonstrates long-context capability well beyond the original LLaMA 2K/4K window through its memory mechanism.
03
Long-context tasks: Shows improvements on passkey retrieval and long-form summarization tasks that stress long-range attention.
04
Efficient extension: Part of the 2023 explosion of context-window-extension techniques that would culminate in ~1M-token proprietary models the following year.