🚀NEW LABGetting Started with Claude AgentsStart lab
Memory · Efficiency

ThinK

First page
ThinK
Paper summary

proposes an approach to address inefficiencies in KV cache memory consumption; it focuses on the long-context scenarios and the inference side of things; it presents a query-dependent KV cache pruning method to minimize attention weight loss while selectively pruning the least significant channels

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack