Can Agents Design Libraries for Agents?

Gabriel Orlanski, Frederic Sala and Aws Albarghouthi (UW-Madison), Alex L. Zhang (MIT), Vincent Sunn Chen (Snorkel AI) and Ludwig Schmidt (Stanford) introduce LibraryDesignBench, which scores a library written by one agent only by how well other agents can use it.
Ask this paper
Design. A designer agent implements a full library from a capability specification that does not prescribe the API. Three consumer agents from different model families then solve 242 expert-validated problems across 15 tasks in four languages, scored on correctness and on program simplicity.
Agents reproduce human abstractions. On 11 of 15 tasks, agent designers arrive at the same abstractions as the human-written production library.
Consumers underuse libraries. Downstream agents adopt agent-written and human-written libraries at similar rates but reimplement capabilities the library already provides. The failure audit attributes most excess code to rigid or hard-to-use interfaces, not to missing features.
Agent-first guidance helps. Having GPT-6 Astra sketch consumer programs first, ship them as runnable examples and test the design with subagents raises the score by 2.3 points, mostly from 6.8% higher simplicity, though it still falls just short of the production library.
Abstract
Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-validated programming problems across 15 library-design tasks in four languages. On eleven of the fifteen tasks, agent designers reproduce the abstractions of the human-written production library. Downstream agents adopt agent- and human-written libraries alike but underuse them, reimplementing capabilities the library already provides. Our failure analysis finds that downstream agents write extra code mainly because agent-written libraries are rigid or hard to use, not because capabilities are missing. We also experiment with giving designers more prescriptive, agent-first guidance and having them test their library with subagents; this improves downstream scores and yields simpler programs. LibraryDesignBench provides both a testbed for evaluating library-design practices for agent users and an initial design baseline that improves downstream reuse.