🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 24, 2026
Agents

Emergent Collusion in Long-Horizon LLM Agent Interaction

First page
Emergent Collusion in Long-Horizon LLM Agent Interaction
The curator’s take

Xinrui Shi and Diyi Yang (Stanford) with Yanzhe Zhang (Georgia Tech) show that two LLM agents that repeatedly verify each other's work drift into collusion when following the verification protocol conflicts with maximizing reward.

Ask this paper

Key points
01

Environment. Two agents complete individual tasks, share logs, verify each other and receive rewards over many rounds, under realistic constraints that make honest verification cost reward.

02

Prevalence. Collusion appears in 94% of trajectories across 10 models.

03

Capability and speed. Within a model family, more capable models reach collusion earlier.

04

Peer effects. Controlled interventions on one agent's behavior change whether the other colludes.

05

Levers. Reward structure, verification feedback and interaction history all matter, and restricting how much history agents can see reduces collusion.

Abstract

LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs, verify each other's work, and receive rewards. We introduce realistic constraints that make compliance with the verification protocol incompatible with reward maximization, and find that agents increasingly deviate from the protocol over repeated interactions. Collusion emerges in 94% of trajectories across 10 models, and more capable models within the same family reach it earlier. Controlled peer interventions show that collusion is shaped by peer behavior, while ablations reveal additional effects of reward structure, the verification feedback agents receive, and their interaction history. In particular, restricting the amount and scope of interaction history available to agents reduces collusion. Overall, our findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks.

Every Monday
Get next week’s papers.
Subscribe on Substack