MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

Ruike Cao, Fanyu Zhao and colleagues at Alibaba's Qwen Applications Business Group, USTC and Fudan introduce MemCalib, a benchmark for whether a model gives each retrieved memory the right amount of influence on its answer, and MemCalib-RL to train that behavior.
Ask this paper
Target capability. Memory systems can retrieve the right items, but the model still has to decide how much each memory should shape the response.
Frontier models miscalibrate. Open and closed frontier models often over-use or under-use memories instead of matching each proposition to its target level.
Standard post-training is skewed. GRPO and on-policy self-distillation improve one direction of error while making the other worse.
MemCalib-RL. An ordered bidirectional counterfactual credit-assignment method separates over-use and under-use signals and assigns credit to response tokens through exact atom ablation.
Results. On Qwen3-8B, Ministral-3-8B-Instruct and Qwen3.5-35B-A3B it gives the best overall score with better balance between the two errors, and gains carry over to external benchmarks.
Abstract
The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a benchmark grounded in realistic memory-system scenarios for evaluating memory use and advancing optimization algorithms. Results on the MemCalib test set reveal that frontier open- and closed-source models struggle to use memory appropriately. They frequently over-use or under-use memory rather than matching each proposition's actual use to its target level, leading to biased, low-quality responses. Experiments with common post-training algorithms, including group relative policy optimization and on-policy self-distillation, further reveal a clear directional skew: trained models improve in one direction while deteriorating in the other. We therefore propose MemCalib-RL, an ordered bidirectional counterfactual credit-assignment algorithm that separates over- and under-use signals and localizes their credit to response tokens through exact atom ablation. Results across model families and scales (Qwen3-8B, Ministral-3-8B-Instruct, and Qwen3.5-35B-A3B) show that MemCalib-RL achieves the best overall performance while better balancing over-use and under-use, with gains generalizing beyond MemCalib in external benchmark evaluation. Further experiments support its design choices and robustness and provide insight into its training dynamics.