🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 7, 2026
Agents

Causal Improvement Graph for Agentic Harness Optimization

First page
Causal Improvement Graph for Agentic Harness Optimization
The curator’s take

Junjie Zhang, Shunyu Liu and Dacheng Tao (Nanyang Technological University) with Ting-En Lin and Yongbin Li (Tongyi Lab, Alibaba) introduce the Causal Improvement Graph (CIG), which stores the state of an automated harness-optimization search in a persistent graph instead of in the proposer's growing history.

Ask this paper

Key points
01

Graph. Evidence, Hypothesis, Intervention and Outcome nodes record what was observed, how it might be explained, how to test that, and what the evaluation showed. Proposers act on local parts of the graph rather than rereading raw history.

02

Main results. With GPT-5 mini as task and optimization agent, CIG reaches 60.0 on SWE-bench Verified, 44.4 on AppWorld and 29.4 on a Terminal-Bench 2.0 subset, a 44.6 macro average against 42.3 for HarnessFix and 39.5 for Meta-Harness.

03

Over the initial harness. Gains are 14.3, 8.8 and 11.8 points, and CIG beats a full-history baseline with the same initialization and evaluation allowance by 4.5 points on average.

04

Robustness. On classification tasks, gains over full history hold across five solver models, from 3.5 to 12.5 points, and across proposer models; ablations support each structural component.

Abstract

Agentic Harness is the runtime that constructs task context and controls execution flow, thereby shaping overall agent performance. Given a fixed model and external evaluation, automated Harness optimization seeks to improve this runtime through an iterative proposal--evaluation loop to better solve target tasks. Existing meta-harness methods mainly adopt proposer-centric discovery, in which an LLM-based proposer integrates accumulated experimental findings to determine subsequent Harness revisions. This places the burden of maintaining the evolving improvement state on the proposer as history expands and its underlying experimental logic becomes harder to discern. In this paper, we introduce the Causal Improvement Graph (CIG), a graph-governed meta-harness framework that externalizes the evolving improvement state in a persistent graph, allowing prior findings to directly govern subsequent Harness optimization through local proposer operations. CIG grows and links Evidence, Hypothesis, Intervention, and Outcome nodes to represent what was observed, how it may be explained, how to test that explanation, and what the evaluation reveals. Their structural relations preserve how the improvement state changes across iterations, allowing local proposers to build directly on relations among prior findings rather than recover them from raw history. Across various agent tasks, CIG discovers stronger Harnesses than previous meta-harness baselines and remains robust to the choice of task solver and proposer. Structural ablations further support the design of an explicit improvement state with graph-governed evolution.

Every Monday
Get next week’s papers.
Subscribe on Substack