Just-In-Time Agent Memory with Runtime Agentic Research

Bingyu Yan, Zheng Liu and colleagues at the Beijing Academy of Artificial Intelligence propose Just-In-Time Agent Memory (JAM), which keeps complete raw histories and builds query-specific context at request time with a trained Researcher agent, instead of compressing memory before requests arrive.
Ask this paper
Design. A Memorizer stores full histories in a hierarchical page store with short navigational summaries. At query time a Researcher retrieves, inspects and combines evidence until it judges the context sufficient.
Training data. Memory-Gym synthesizes evidence-grounded training scenarios across nine task types and six domains, covering single-hop lookup, multi-hop evidence linking and multi-session summarization.
Training. Verified-trajectory SFT followed by hint-guided GRPO raises the Researcher's overall accuracy from 54.08% untrained to 65.67% after SFT and 75.54% after RL.
Results. With a Qwen3.5-4B backbone, JAM beats the strongest baseline in its group by 12.54, 20.00 and 11.20 points on LoCoMo, NarrativeQA and HotpotQA, and outperforms 7B to 14B trained memory agents in most settings.
Latency. Against the trained MemAgent baseline, query latency drops from 58.08 s to 13.81 s while F1 rises from 46.13 to 52.09. Median research rounds are 3 on most benchmarks.
Abstract
Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine-grained information that later becomes important. To address this limitation, we propose Just-In-Time Agent Memory (JAM), a trainable framework for query-conditioned context construction at runtime. A Memorizer preserves complete raw histories in a hierarchical page-store with compact navigational summaries, while a Researcher iteratively retrieves, inspects, and integrates evidence for each request. To train these memory-use behaviors, we introduce Memory-Gym, an evidence-grounded data synthesis pipeline covering nine task types across six domains, and optimize the Researcher through verified-trajectory supervised fine-tuning followed by Hint-guided Group Relative Policy Optimization. We demonstrate the effectiveness of JAM across a variety of benchmarks on agent memory and long-context processing, where it achieves stronger task performance than AOT-style memory systems while remaining substantially more efficient than prior trained agentic memory approaches. To support reproducibility and future research, we release our anonymized source code at https://github.com/VectorSpaceLab/general-agentic-memory.