🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 9, 2026
Agents

MIMESIS: Learning User Simulators as Training Environments for Interactive Agents

First page
MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
The curator’s take

Hoang Phan, Dat Huynh, Deren Lei and colleagues at Meta Superintelligence Labs (with NYU) build MIMESIS, a trained user simulator meant to serve as the training environment for interactive agents.

Ask this paper

Key points
01

Problem. Off-the-shelf assistant LLMs playing users are too cooperative, explicit and uniform. With a fixed GPT-5.5 agent, frontier-API users make tau-bench tasks noticeably easier than real users do.

02

Simulator. A 9B model trained on human conversations with explicit reasoning supervision and 13 behavioral patterns mined from real users. It scores 65.7 on SOUL-Index, above the strongest frontier model.

03

Fidelity. Against Claude Opus 5, the strongest baseline, MIMESIS improves behavioral fidelity by 13.4 points on RealUserSim and cuts Turing distance by 3.6 points on SimulatorArena.

04

Agent training. Agents trained by multi-turn RL against the frozen simulator across eight environments outperform agents trained against GPT-5.5 under all nine unseen user simulators.

05

CSD. Coached On-Policy Self-Distillation turns the simulator's private reasoning and next utterances into coaching notes and dense token-level supervision, adding gains across all nine evaluation users.

Abstract

Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale. Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents. In contrast, most agent-training frameworks rely on off-the-shelf assistant LLMs, whose helpfulness can make them overly cooperative, explicit, and behaviorally homogeneous compared with real users. We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions. Empirically, our 9B model achieves a SOUL-Index of 65.7, surpassing the strongest frontier model. Compared with Claude-Opus-5, the strongest baseline on RealUserSim and SimulatorArena, MIMESIS improves behavioral fidelity by 13.4 points and reduces Turing distance by 3.6 points, respectively. We then freeze the simulator and train an agent by interacting with the frozen simulator using multi-turn reinforcement learning. Across eight environments, training with MIMESIS yields better agent performance than training with GPT-5.5 under all nine unseen user simulators, demonstrating stronger generalization to new user simulators. Moreover, we propose Coached On-Policy Self-Distillation (CSD), which leverages simulator-generated private reasoning traces and subsequent utterances as feedback on how well the agent addresses user needs. A coach converts this information into concise coaching notes that describe how the agent can better anticipate user needs and adapt its behavior over the course of an interaction. CSD turns this feedback into dense, token-level supervision beyond sparse task rewards, yielding further gains across all nine evaluation user models.

Every Monday
Get next week’s papers.
Subscribe on Substack