🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 3, 2026
Agents

Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

First page
Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams
The curator’s take

Sahan Paliskara (independent), Nattaput Namchittai (Stanford), Andrew Lampinen (Anthropic) and colleagues study what happens when several agents, each acting for a different user, share a resource such as a compute budget, a calendar or a release cutoff.

Ask this paper

Key points
01

Setup. 77 scenarios in four environments (shared API-key budget, clinic calendar, group order or booking, merge queue) and five frontier models, comparing one coordinator agent serving all users with teams of one agent per user, with and without a communication channel.

02

Teams lose everywhere. Teams deliver worse group outcomes than the coordinator in every environment; without a channel they collapse completely in two environments, and in the personal-assistant environment the coordinator fulfills a targeted request about twice as often as teams.

03

Failure behaviors. Teams stall as they grow, override each other's actions and fabricate claims.

04

Mitigations. A team lead, explicit procedural instructions, and a platform check that forces agents to read peers' messages before committing all help, but each works only in some environments. Three environments are released as MAMUBench (74 scenarios).

Abstract

People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user with different goals, coordination often fails, and the group ends up worse off than if a single agent had acted for everyone. We study this multi-user, multi-agent setting across five frontier models and 77 scenarios in four environments: an API key environment in which agents share a compute budget, a clinic in which they share a calendar, a personal assistant environment in which they share a group order or booking, and a merge queue in which they share a release cutoff. In each scenario, we compare a single agent that serves every user (a coordinator) to a team in which each agent serves one user, with and without a communication channel between the agents. Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps. For example, in the personal assistant environment, the coordinator fulfills a targeted user request about twice as often as teams. We identify distinct behaviors associated with this poor group-level performance, including stalling as teams grow, overriding each other's actions, and fabricating claims. We find effective but environment-specific mitigations, such as a team lead, explicit procedural instructions, and a platform check that makes an agent read its peers' messages before committing. We will release the API key, clinic, and personal assistant environments as MAMUBench, comprising 74 scenarios for evaluating multi-user, multi-agent coordination.

Every Monday
Get next week’s papers.
Subscribe on Substack