Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild

Yifan Xiong and Yiling Lou (UIUC) with Jingyi Ge (UC Berkeley) and Zhenpeng Chen (Tsinghua) run the first empirical study of test adequacy in agent harnesses and introduce HarnessTester, a test generator built for harness code.
Ask this paper
Study. Across 10 widely used agentic systems including OpenClaw, average line coverage is below 50% and branch coverage below 40%. LLM-dependent harness code, which parses model outputs and controls agent behavior, has less than half its lines and branches covered and a 33.60% mutation score.
HarnessTester. It encodes explicit agent-harness contracts so generated tests set up realistic model outputs and exercise the LLM-dependent paths.
Coverage. Against the strongest general-purpose test generator per agent, it delivers 75.95% larger line-coverage gains, 84.76% larger branch-coverage gains and 69.89% larger mutation-score gains.
Bugs found. On 250 historical harness bugs its tests detect 43 against 2 for the best baseline. On current versions it found 122 harness bugs, 88 previously unknown and 69 confirmed by developers.
Abstract
LLM-based agentic systems are emerging as a new software paradigm. Modern agents are typically composed of backbone LLMs and a surrounding harness that serves as the operational software infrastructure for agent execution. As agent harnesses grow increasingly complex, agents suffer from diverse harness implementation bugs, raising substantial reliability concerns. In this work, we conduct the first empirical study to systematically investigate the test adequacy of harness in real-world agentic systems. Our analysis reveals that agent harness remains substantially undertested. In particular, LLM-dependent harness (LDH) code, despite its critical role in processing LLM outputs and governing agent behavior, receives limited testing attention, with less than half of its lines and branches covered by existing tests. Motivated by these findings, we further propose HarnessTester, the first harness-oriented test generation technique that incorporates explicit agent-harness contract support to construct contract-faithful test setups and extensively exercise LDH code. Our evaluation shows that HarnessTester substantially outperforms state-of-the-art general-purpose test generation techniques in achieving 75.95%/84.76% larger line/branch coverage gains and 69.89% larger mutation-score gains. Furthermore, HarnessTester detects 122 real-world harness bugs in widely-used agentic systems (e.g., OpenClaw), among which, 88 bugs are previously-unknown bugs and 69 bugs have been confirmed by agent developers. These results highlight the practical effectiveness of HarnessTester in improving test adequacy and assuring the reliability of real-world agentic systems.