🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 22, 2026
Agents

A Study of the Reliability of Agentic AI-Generated Programs

First page
A Study of the Reliability of Agentic AI-Generated Programs
The curator’s take

Ayesha Shafique, Barton P. Miller and Elisa R. Heymann rebuild ten release-quality Linux utilities, including dash, make, grep, less and tnftp, with a best-practices Claude Code workflow and fuzz both the AI versions and the human originals to compare reliability.

Ask this paper

Key points
01

Method. The AI versions were specified, designed, coded and tested with Claude Code on Opus 4.8, then tested with classic black-box generational fuzzing and coverage-guided mutational fuzzing with AFL++.

02

Headline. The AI-generated utilities were typically as reliable as the latest human-written releases, and often more reliable, with fewer failures overall.

03

Different failure profile. The AI code had fewer memory errors such as buffer overflows but more hangs such as infinite loops.

04

Supervision still required. Code quality depended heavily on the prompts and skills used and on how the human directing the process responded to the agent.

05

Workflow as specification. The authors argue that the agentic workflow, with its prompts and skills, works as a specification of the code that makes later maintenance cheaper.

Abstract

Agentic-AI based software development offers the promise of faster completion of the software, greater programmer efficiency, and more reliable code. The question is how can we verify these claims in an objective way? In this project, we attempted to answer this question based on three practices. First, we applied a typical best-practices agentic AI workflow for software development. Second, our target programs were ten well-known, release-quality human-written Linux utility programs so that we could compare the AI-generated code against a concrete ground truth. Third, we based our measure of reliability on a widely used testing technique, fuzz random testing. For this testing, we used both classic black box, generational testing and more modern coverage guided (gray box, mutational) testing using AFL++. We found that the AI-generated versions of the utility programs were typically as reliable - often more reliable - than the latest human-generated versions of these programs. While the AI-generated versions did have some failures, they were less common than the code from the standard repositories. Interestingly, the AI-generated code was less likely to have failures such as memory errors (such as buffer overflows) but more likely to have hangs such as infinite loops. In addition, we verified that generating robust and reliable software using agentic AI requires careful practice and human supervision. The quality of the code is highly dependent on the prompts and skills used, and how the human directing the process responds. We also demonstrated that using agentic AI workflow for software development (with its prompts and skills) can become a specification of the code that leads to cost-effective sustainability of the software.

Every Monday
Get next week’s papers.
Subscribe on Substack