🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 8, 2026
Code

Harness Engineering for Software Engineering via Modular Executable Dev-Primitives

First page
Harness Engineering for Software Engineering via Modular Executable Dev-Primitives
The curator’s take

Haibo Jin, Xinjie Li, Peng Kuang and Haohan Wang at the University of Illinois Urbana-Champaign propose Dev-Primitives, which pair each repository artifact with a resident LLM, and HERMES, a harness that activates these primitives at repository scale.

Ask this paper

Key points
01

Problem. Terminal-based coding agents must repeatedly rebuild program state spread across files, configs, tests and dependencies, which lengthens histories and causes context explosion on long workflows.

02

Dev-Primitives. Each artifact gets an agent interface grounded in its own implementation and dependencies, so components can reason in natural language, message each other and modify themselves locally.

03

HERMES. Dependency-aware dynamic activation picks which primitives to wake for a task, and a bug-diagnosis step maps execution evidence back to the components that must change.

04

Results. Across four software engineering benchmarks HERMES beats matched baseline harnesses by 12.4 points on average; with GPT-5.6 Sol on one benchmark the composite score rises from 6.5% to 31.0% at medium effort.

05

Small models. With strong activation and diagnosis models, Qwen3-8B Dev-Primitives stay within 4.5 points of an all-GPT-5.6 Sol configuration and cut Terminal-Bench 4.0 inference cost by 26.2%.

Abstract

Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across source files, configurations, tests, dependencies, and runtime behavior, leading to increasingly long interaction histories, context explosion, and semantic drift. Large repositories further complicate the identification of task-relevant components. To address these challenges, we introduce \textbf{Dev-Primitives} (\emph{Development Primitives}), a modular and executable abstraction that transforms repository components from passive software artifacts into active participants in software engineering. Each Dev-Primitive pairs a repository artifact with a resident LLM, which gives the artifact an agent-native interface grounded in its own implementation and dependencies, enabling natural-language reasoning, inter-component communication, and localized self-modification. Building on Dev-Primitives, we propose \textbf{HERMES}, a Harness Engineering framework for software engineeRing via Modular Executable Dev-PrimitiveS, which instantiates these primitives at repository scale through a dependency-aware dynamic activation mechanism and a bug diagnosis mechanism that maps execution evidence back to the components that must be revised. Extensive experiments on four software engineering benchmarks demonstrate that HERMES outperforms matched baseline harnesses by 12.4\% on average. Moreover, when paired with strong activation and diagnosis models, HERMES, even with Qwen3-8B Dev-Primitives, remains within 4.5\% of the homogeneous GPT-5.6 Sol configuration across all four benchmarks, while reducing inference cost by 26.2\% on Terminal-Bench 4.0, highlighting the importance of harness design in software engineering agents.

Every Monday
Get next week’s papers.
Subscribe on Substack