🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 24, 2026
Memory · Training · Agents

PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents

First page
PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents
The curator’s take

Pirzada Suhail, Menglin Xia and colleagues at Microsoft Research and M365 propose Pseudo Self-Distillation (PSD), which trains small Qwen3 models to run a multi-stage memory-construction pipeline that normally needs GPT-4.1-mini, using only the oracle's text outputs.

Ask this paper

Key points
01

One model, two roles. The same small model acts as teacher when its prompt contains the oracle's answer as privileged context and as student when it sees only the task; the student learns to match the teacher's output distribution.

02

No logits needed. Oracle knowledge enters only through the prompt, so a closed API model can serve as the source without exposing hidden states.

03

Target pipeline. PSD replaces GPT-4.1-mini in every stage of MEMORA (segmentation, episodic and factual memory extraction, consolidation, cue anchors), after an SFT stage that fixes JSON and schema failures.

04

LoCoMo. A PSD-trained Qwen3-4B scores 0.8699 overall LLM-judge, above MEMORA with GPT-4.1-mini (0.849) and full context (0.825); at 1.7B it already beats every non-MEMORA baseline, and 0.6B matches 1.7B at 0.8211.

05

Transfer. Trained only on LoCoMo, the students keep strong memory-construction quality on a 100-question LongMemEval subset.

Abstract

Memory systems are becoming a core component of LLM agents, but constructing and maintaining memory remains expensive because it relies on repeated calls to large proprietary language models. This cost creates a major barrier to deploying memory-enhanced agents at scale. In this paper, we present Pseudo Self-Distillation (PSD), a framework that enables small language models (SLMs) to construct hierarchical memory representations by distilling behavior from a strong black-box oracle through a multi-stage training pipeline. Standard distillation methods require access to teacher logits or hidden states, which closed models do not expose. Unlike conventional self-distillation settings, where supervision is derived from a model's own predictions, sampled rollouts, or aggregated outputs, PSD enables a single-model distillation setup while channeling external oracle knowledge through the prompt. PSD uses a single small model in two roles: a teacher that sees a privileged prompt containing the oracle's answer as reference context, and a student that sees only the task prompt. The student learns to reproduce the teacher's output distribution, absorbing oracle-guided behavior into its own weights without accessing the oracle's internals. On LoCoMo, PSD-trained Qwen3-0.6B, 1.7B, and 4B match or exceed GPT-4.1-mini on downstream retrieval at a fraction of the deployment cost, with off-policy PSD achieving the strongest results across most conditions. We further show that this memory-construction capability transfers out-of-distribution to LongMemEval, despite the students being trained exclusively on LoCoMo with no exposure to LongMemEval data.

Every Monday
Get next week’s papers.
Subscribe on Substack