🚀NEW LABGetting Started with Claude AgentsStart lab
Efficiency

Learning to Compress Prompts with Gist Tokens

First page
Learning to Compress Prompts with Gist Tokens
Paper summary

Trains LMs to compress prompts into reusable "gist" tokens.

Ask this paper

Key points
01

Prompt compression: Compresses long prompts into a small set of gist tokens that encode the same instruction information.

02

26x compression: Achieves 26x prompt compression with negligible quality loss on downstream tasks.

03

Up to 40% FLOPs reduction: Substantial inference-time compute savings on repeated prompts.

04

Production optimization: Particularly valuable for systems with long system prompts reused across many requests - a pattern that became ubiquitous in 2024 agent systems.

Every Monday
Get next week’s papers.
Subscribe on Substack