Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems

Zhongwen Luan, Xiaoyu Zhang, Ming Hu and coauthors ask whether multi-agent repair methods causally fix failures or merely exploit LLM sampling randomness, and build SymTrace to make the distinction measurable.
Ask this paper
Deterministic replay with intervention anchors: SymTrace reconstructs execution before an anchor from recorded logs and regenerates only the downstream trajectory, which is what makes failure reproduction reliable.
SymFail dataset: 536 human-annotated failure trajectories with graph-linked locations, categories and trace evidence, across three mainstream multi-agent frameworks.
Unguided rerun is close to useless: Only 67.97% failure reproduction and 6.90% repair rate, which reframes most reported multi-agent self-repair as resampling luck.
Symptom-driven intervention is 191.89% better: 20.15% of failed cases repaired, a nearly 3x relative improvement over state-of-the-art repair methods, though the absolute number shows how open this problem still is.
Abstract
As large language model (LLM)-based multi-agent systems (MASs) are increasingly applied to long-horizon complex tasks, their reliability has emerged as the core bottleneck hindering their real-world deployment. Existing MAS debugging and repair methods typically rely on rerunning and resampling the entire execution trajectory. However, a fundamental question remains to be answered: do these methods causally repair MAS failures or merely stochastically repair by leveraging the randomness of LLM sampling? To evaluate the effectiveness of MAS repair methods, we introduce SymTrace, a controlled evaluation framework that records the MAS execution trajectory and establishes intervention anchors. During replay, it effectively reconstructs the execution before the anchor using recorded logs and only regenerates the downstream trajectory, thereby enabling the reliable reproduction of MAS failures. We further construct the dataset SymFail, comprising 536 human-annotated failure trajectories with graph-linked locations, categories, and trace evidence. Based on these foundations, we conduct a large-scale empirical study across three mainstream MAS frameworks. Our findings reveal that existing unguided rerun methods are highly unreliable, exhibiting low failure reproduction and repair rates (only 67.97% and 6.90%, respectively). Building upon these findings, we further explore the effectiveness of a symptom-driven intervention method, which successfully repairs 20.15% of the failed cases (a 191.89% improvement to state-of-the-art repair methods). This study aims to provide actionable insights for MAS debugging and repair research, paving the way for the robust deployment of multi-agent systems.