AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

Raphael Shu (OpenAgents), Yusen Zhang (Columbia), Young Min Cho (Penn) and colleagues (COLM 2026) introduce AgentWorld, a benchmark for long-horizon collaboration among 3 to 20 LLM agents with asymmetric roles in an MMORPG sandbox.
Ask this paper
Tasks. 100 human-annotated tasks plus 100 augmented variants, each spanning 50+ interaction rounds. Agents act independently without access to each other's internal state and must coordinate through messages, joint plans and shared resources.
New metric. Causal Collaboration Effectiveness (CCE) traces causal dependencies between actions and measures the fraction of a team's actions that actually contributed to the outcome.
Results. Gemini 3 Flash leads at 52.0% task success, followed by Claude Haiku 4.5 (45.0%), GPT-5 Mini (36.0%) and DeepSeek R1-70B (20.0%). The best CCE is 0.320, so fewer than a third of actions advance the goal.
Hardest categories. Survival tasks are easiest (67-100% success) while coordination (12%) and construction (20%) are hardest.
Failure modes. Communication breakdowns, role confusion and failure to keep shared plans across rounds.
Abstract
Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with 100 augmented variants) for evaluating long-horizon, multi-agent collaboration. Tasks span 50+ interaction rounds across a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others' internal states. To quantify collaboration effectiveness in addition to conventional binary task success, we propose Causal Collaboration Effectiveness (CCE), a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B show that even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. AgentWorld is fully open-source.