How does Adversarial Influence Scale in Multi-Agent Systems?

Addison J. Wu, Jasin Cekinmez, Michel Liao, Karthik Narasimhan and Thomas L. Griffiths (Princeton University) run multi-agent deliberation on Humanity's Last Exam questions with groups of 2 to 21 agents and a varying number of instructed deceivers, and find that the fraction of deceivers, not the group size, determines how often honest agents give up correct answers.
Ask this paper
Linear in proportion. The defection rate, how often an initially correct agent switches to a wrong final answer, rises roughly linearly with the fraction of deceivers for Gemini, Grok, DeepSeek and Muse Glimmer. At a fixed fraction, larger groups do not show consistently lower defection.
Minority deceivers work. Human conformity studies show reliable influence only when confederates are a majority; LLM agents defect regularly when deceivers are a minority.
Coordination hurts deceivers. A private planning channel lowered defection from 29.2% to 21.9% for Gemini 3.8 Flash and from 30.0% to 24.6% for Grok 4.3. Honest agents voiced suspicion of coordination in only 1 of 60 coordinated trials.
Model choice. The honest model matters more than the deceiver. The more sycophantic Muse Glimmer defected at 37.7% versus 19.5% for Gemini; the more persuasive DeepSeek caused 26.8% defection versus 21.2% for Grok. Most first defections happen in rounds 1 and 2.
Implication. Adding more agents does not defend a deliberation system, because an adversary can add deceivers in proportion to the group.
Abstract
Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent. It is not the number of agents in the group that matters, but the proportion of deceivers. We observe that the defection rate, how often initially correct agents switch to an incorrect final answer, rises linearly with this proportion. Whereas humans in comparable conformity studies are reliably swayed only when misleading confederates form a majority, LLM agents defect regularly even when deceivers remain a minority. Susceptibility also depends on which models are interacting, especially on the honest agent side. Unexpectedly, allowing deceivers to coordinate privately can make them less effective. Altogether, our results show that adding more agents is therefore not a sufficient defense, because the adversary can simply scale with the group.