🚀NEW LABGetting Started with Claude AgentsStart lab
Memory · Retrieval · Evaluation

Diagnosing Agent Memory

First page
Diagnosing Agent Memory
Paper summary

This paper introduces a diagnostic framework that separates retrieval failures from utilization failures in LLM agent memory systems. Through a 3x3 factorial study crossing three write strategies with three retrieval methods, the authors find that retrieval is the dominant bottleneck, accounting for 11-46% of errors, while utilization failures remain stable at 4-8% regardless of configuration. Hybrid reranking cuts retrieval failures roughly in half, delivering larger gains than any write strategy optimization.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack