Kernel-Managed Shared Memory for System-Wide Personalization

Ryan Lum and Yongfeng Zhang (Rutgers University) move memory retrieval, privacy enforcement and prompt injection out of individual agents and into the agent-system kernel, and evaluate the design on AIOS across three assistant models and 1,800 trials.
Ask this paper
Kernel owns retrieval and injection: Specialized agents write structured tagged memories, but the kernel decides what is retrieved, what privacy rules apply, and what goes into the prompt. Context learned by one agent therefore becomes available to others without each agent implementing its own policy.
Against an unmanaged backend: Compared with Mem0 on identical underlying storage, kernel-managed retrieval improves personalization scores by 2.4 to 4.0 points on a 5-point scale (profile usage moves from 1.05 to 4.69 on GPT-4o), with every comparison significant at p below 10^-18.
Against full context concatenation: Unfiltered concatenation is the ceiling on available context. Kernel-managed injection matches it statistically on two of three models and shows a small deficit on the third.
Cost side of that match: End-to-end latency is 15 to 61% lower across all three models, with matching reductions in per-call tokens and inference cost, because the prompts are substantially shorter.
Scope: Three assistant models (GPT-4o, Llama-3.1 8B, Qwen-2.5 7B) across 1,800 total trials, with four memory designs compared.
Abstract
AI systems become more useful when they can adapt to the people using them, but in multi-agent systems, useful context learned by one agent often remains unavailable to others. We present kernel-managed shared memory, a system-level abstraction in which specialized agents write structured, tagged memories while the agent-system kernel, not individual agents, governs retrieval, privacy enforcement, and prompt injection. We implement and evaluate this design on AIOS and compare it against three alternatives across three assistant models (GPT-4o, Llama-3.1:8B, Qwen-2.5:7B) and 1,800 total trials. Against an unmanaged external memory backend (Mem0) using identical underlying storage, kernel-managed retrieval and injection improve personalization scores by 2.4-4.0 points on a 5-point scale (e.g., 1.05 to 4.69 profile usage on GPT-4o), with every comparison significant at p < 10^-18. Against standard retrieval-augmented injection, gains are similarly large and consistent across all three models. Against full, unfiltered context concatenation, a soft ceiling on available context rather than on response quality, kernel-managed injection statistically matches performance on two of three models and shows a small, model-specific deficit on the third, while using substantially shorter prompts: end-to-end latency is 15-61% lower across all three models, with corresponding reductions in per-call token usage and inference cost. These results indicate that centralizing memory management in the agent-system kernel, rather than leaving retrieval and privacy enforcement to individual agents, delivers most of the personalization benefit of unconstrained context at a fraction of its cost.