🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 1, 2026
Agents

AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems

First page
AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems
The curator’s take

Yiming Cheng, Zhenpeng Chen, Yiling Lou and colleagues (Fudan, UChicago, Tsinghua) present AgentBug-Smith, which continuously finds and reproduces real bugs in the harnesses of open-source agentic systems, and use it to build Live-Harness-Bench.

Ask this paper

Key points
01

Reproduction. Across backbone LLMs, AgentBug-Smith reproduces harness bugs at 10.67% to 27.56% higher success rates than bug-reproduction methods built for general software.

02

Live benchmark. Live-Harness-Bench currently holds 200 reproducible harness bugs and grows as new bugs are mined, where existing benchmarks are small, fixed and took hundreds of human hours to build.

03

Agents struggle. State-of-the-art software agents repair few real-world harness bugs on the benchmark.

04

Reusable fixes. Repair skills distilled from the benchmark's fixes raise harness-bug repair rates by 6.32%.

Abstract

Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number of executable harness bugs while requiring hundreds of human hours to construct. This work presents AgentBug-Smith, an automated harness bug reproduction approach that continuously discovers and reproduces real-world harness bugs from open-source agentic systems. Across different backbone LLMs, AgentBug-Smith consistently outperforms existing bug reproduction techniques designed for general software, achieving 10.67% - 27.56% higher success rates of reproducing harness bugs. By applying AgentBug-Smith to open-source agentic systems in the wild, we construct Live-Harness-Bench, a live and extensible benchmark that currently contains 200 reproducible harness bugs. We further demonstrate the utility of Live-Harness-Bench through two downstream applications. First, we use Live-Harness-Bench as the evaluation benchmark to systematically evaluate state-of-the-art software agents, revealing their limited capabilities in repairing real-world harness bugs. Second, we use Live-Harness-Bench as a knowledge base of real-world harness bug fixes, from which reusable repair skills can be distilled to improve existing software agents, increasing their harness-bug repair rates by 6.32%. Together, AgentBug-Smith and Live-Harness-Bench establish a scalable foundation for continuously evaluating and improving software agents on harness bug repair, turning real-world agent failures into executable evaluation instances and reusable knowledge for harness improvement, thus contributing to the ultimate goal of recursively self-improving agents.

Every Monday
Get next week’s papers.
Subscribe on Substack