Learning to Compress Prompts with Gist Tokens
First page

Paper summary
Trains LMs to compress prompts into reusable "gist" tokens.
Ask this paper
01
Prompt compression: Compresses long prompts into a small set of gist tokens that encode the same instruction information.
02
26x compression: Achieves 26x prompt compression with negligible quality loss on downstream tasks.
03
Up to 40% FLOPs reduction: Substantial inference-time compute savings on repeated prompts.
04
Production optimization: Particularly valuable for systems with long system prompts reused across many requests - a pattern that became ubiquitous in 2024 agent systems.