When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration

Yaxin Gong, Xiangnan He and colleagues at USTC and the Qwen Business Unit of Alibaba (with NUS) run controlled experiments that separate the benefit of inter-agent messages from the damage done by wrong ones.
Ask this paper
Design. Across five benchmarks and five receiver models, the downstream task and evidence are fixed and the receiver answers with no message, the upstream agent's original message, or a message with the opposite conclusion.
Messages help. When the receiver would otherwise be wrong, upstream messages often correct it.
Messages hurt. When the receiver would be right on its own, a wrong upstream message flips the answer in up to 32% of cases.
Answer substitution. In 94% of audited harmful cases the receiver copies the upstream agent's specific wrong answer rather than making a new error. Dropping unreliable messages recovers part of the lost accuracy, which argues for gating messages on upstream reliability and the receiver's own evidence.
Abstract
Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its own evidence. Prior work has not clearly separated the benefits of communication from the damage caused by incorrect messages. We study this problem with controlled experiments across five benchmarks and five receivers, keeping the downstream task and evidence fixed while comparing answers under three conditions: no message, the upstream agent's original message, or a message with the opposite conclusion. Our experiments reveal three key findings. First, messages often help when the downstream agent would otherwise answer incorrectly. Second, messages can also hurt: when the downstream agent would answer correctly without a message, an incorrect upstream message changes the answer in up to 32% of cases. Third, in 94% of audited harmful cases, the downstream agent copies the upstream's specific wrong answer--a pattern we term answer substitution. Removing unreliable messages recovers part of the lost accuracy, suggesting that communication should be selective based on upstream reliability and the evidence already available to the downstream agent.