Learning to Compress Prompts with Gist Tokens
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsTrains LMs to compress prompts into reusable "gist" tokens.
01
Prompt compression: Compresses long prompts into a small set of gist tokens that encode the same instruction information.
02
26x compression: Achieves 26x prompt compression with negligible quality loss on downstream tasks.
03
Up to 40% FLOPs reduction: Substantial inference-time compute savings on repeated prompts.
04
Production optimization: Particularly valuable for systems with long system prompts reused across many requests - a pattern that became ubiquitous in 2024 agent systems.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack