ActKV: Efficient LLM Agents through Action-Guided KV Cache Management

Zihan Wang, Xuehai Zhou and colleagues at the University of Science and Technology of China introduce ActKV, a KV cache compression method for agent inference that keeps the cache entries that matter most for generating actions.
Ask this paper
Action-weighted criterion. Existing compressors treat all output tokens as equally important; ActKV scores KV entries by how much they contribute to action tokens, since actions drive task progress.
Three mechanisms. Action-oriented eviction based on stable action attention patterns, confidence-driven budget allocation that grows or shrinks the cache with model uncertainty, and page-aware compression primitives with custom kernels.
Accuracy and memory. On long-trace agent tasks it keeps 98.53% of full-cache accuracy while using 25.98% of peak KV memory.
Throughput. 3.97x the token throughput and 3.58x the task throughput of the uncompressed baseline.
Scope. The paper positions KV compression as complementary to context engineering: one decides which text enters the prompt, the other which runtime states are kept.
Abstract
Agentic LLM inference accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput. Existing compression methods emphasize overall output quality, overlooking the asymmetric importance of actions in driving task progress. Our key idea is to establish a compression criterion that values KV entries by their contribution to action generation and prioritizes action quality. However, iterative execution, dynamic memory demands, and scattered action-critical entries pose challenges to eviction policies, budget allocation, and paged memory integration. To this end, we propose ActKV, the first KV cache compression framework tailored for agentic LLM inference. (i) Action-oriented KV cache eviction exploits stable action access patterns to retain entries critical to future actions, supporting reliable task progress under compression. (ii) Confidence-driven adaptive budget allocation uses LLM's intrinsic confidence to adapt the budget to evolving action-critical memory demands. (iii) Page-aware compression management standardizes compression into three primitives with customized kernels, realizing practical throughput gains. On long-trace tasks, ActKV retains an average of 98.53% of FullKV's accuracy with only 25.98% of its peak KV cache memory. It also achieves 3.97 times and 3.58 times FullKV's token and task throughput, delivering state-of-the-art performance.