LongLLaMA
Free while signed in. Answers cite the passages they came from.

Extends LLaMA's context length using a contrastive training process that reshapes the (key, value) space.
Focused Transformer: Uses contrastive training to make memory-augmented attention more discriminative, reducing distraction from irrelevant context.
Length extrapolation: Demonstrates long-context capability well beyond the original LLaMA 2K/4K window through its memory mechanism.
Long-context tasks: Shows improvements on passkey retrieval and long-form summarization tasks that stress long-range attention.
Efficient extension: Part of the 2023 explosion of context-window-extension techniques that would culminate in ~1M-token proprietary models the following year.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack