Mental-Models for Multi-Agent Systems

Hanan Gani, Lulu Shao and Manmohan Chandraker (UC San Diego) equip agents with a learned latent mental model of their counterpart and use it as a decision variable for action selection. Accepted at NeurIPS 2026.
Ask this paper
Representation. An amortized recursive Theory-of-Mind encoder infers first- and second-order mental states (what the partner believes, and what the partner believes about the agent) from the observed interaction history.
Reward model. A belief-conditioned reward model is trained jointly with the mental-state encoder and scores candidate actions relative to the inferred partner state.
Distillation. A policy is then trained with GRPO under this belief-aware signal, so the deployed agent acts on its own without running the explicit partner model at inference time.
Evaluation. The same framework is tested on BigToM, Sotopia and the multimodal MMRole, with augmented ToM supervision, rationales and hard negatives, plus zero-shot transfer to multi-party FANToM. Explicit mental-state modeling improves interaction quality and ToM accuracy over the base agents, including with Qwen2.5-7B on Sotopia.
Abstract
Large foundation models have accelerated progress toward general-purpose agents that interact with humans and other agents through language and multimodal signals. However, robust multi-agent decision-making requires reasoning about what other agents know, intend, and are likely to do under partial observability. Current agentic systems often operate through prompt design, memory, or end-to-end behavioral shaping, but typically do not learn an explicit partner-state representation that can be reused as a decision variable across tasks. We introduce \emph{mental-model-enabled agents}, a framework that equips an agent with a latent mental model of its counterpart, allowing it to infer hidden beliefs, intentions, and likely reactions from the observed history and use these inferences to guide action selection. Our method learns an amortized recursive Theory-of-Mind representation, with first- and second-order mental-state structure, jointly with a belief-conditioned reward model that evaluates candidate actions relative to the inferred partner state. A policy is then learned under this belief-aware signal, yielding an agent that can act independently at inference time while retaining the benefits of explicit partner modeling. We evaluate the same framework on both language-only and multimodal benchmarks. Across these settings, explicit mental-state modeling consistently improves interaction quality and Theory-of-Mind performance over base agentic systems, showing that structured partner modeling is a useful inductive bias for general multi-agent systems. Our code is publicly available at this https URL