🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Memory

You Only Cache Once

Free while signed in. Answers cite the passages they came from.

First page
You Only Cache Once
The curator’s take

a decoder-decoder LLM architecture that only caches key-value pairs once; it involves a cross-decoder stacked upon a self-decoder which efficiently encodes global key-value caches and the cross-encoder reuses the cache via cross-attention; this leads to a significant reduction in GPU memory use without sacrificing capabilities; achieves comparable performance to Transformer in various settings of scaling up model size and number of training token.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack