🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 28, 2026
Agents

Self-Healing Harness for Runtime Oversight of Agent Self-Modification

First page
Self-Healing Harness for Runtime Oversight of Agent Self-Modification
The curator’s take

Sina Tayebati, Amit Ranjan Trivedi and colleagues at the University of Illinois Chicago, with Ranganath Krishnan of Capital One AI Labs, treat agent self-modification as an admission-control problem: the agent may propose changes to its own instructions, but an external runtime gate decides which changes persist.

Ask this paper

Key points
01

Detect, Notice, Heal, Validate. A model-agnostic harness wraps an unmodified agent. The agent writes candidate behavioral rules into an external workspace, where they get provisional authority during evaluation.

02

Admission rule. A rule becomes persistent only after it improves the triggering failure without regressing protected cases beyond a fixed margin. Replay supplies matched evidence; forward trials are a weaker fallback; a corpus-level guard re-tests the accumulated rule set.

03

Collateral regressions are common. Across 16 paired runs on AppWorld, Terminal-Bench and tau2-Bench, the gate rejected 383 replay-decided proposals, and 211 of them (55%) fixed their triggering failure while breaking a case that previously worked.

04

Results. Task completion is higher with the harness in all 16 pairs; repeated-trial reliability is higher in 12, tied in 4, lower in none.

05

Inspectable changes. Because adaptation edits context and leaves weights fixed, admitted changes can be read, reverted and used with closed-weight models.

Abstract

LLM agents can change their own future behavior, raising a basic control question of which self-generated changes should be allowed to persist. We formulate this as admission control for self-modification. The agent may propose changes to its operating instructions, while an external runtime gate controls persistence. We implement this principle as a model-agnostic self-healing harness that runs a Detect, Notice, Heal, Validate loop around an otherwise unmodified agent. The agent authors candidate behavioral rules in an external workspace, where they receive provisional execution authority during evaluation and acquire persistent cross-episode authority only after measured improvement on the triggering failure without regression beyond a fixed margin on protected cases. Replay provides matched evidence when available, forward trials provide a weaker fallback, and a corpus-level guard re-tests the accumulated active rule set. Across 16 matched Baseline and Harness runs spanning AppWorld, Terminal-Bench, and $τ^2$-Bench, the gate rejected 383 replay-decided proposals. Of these, 211 (55%) improved their triggering failure while degrading a case that previously worked. This shows that locally beneficial self-modifications can introduce collateral regressions often enough to materially affect gate decisions, providing direct empirical motivation for external admission control. Task-completion score is higher under the Harness in all 16 pairs, with two paired bootstrap intervals excluding zero, while repeated-trial reliability is higher in 12 pairs, tied in 4, and lower in none. Because adaptation modifies the policy-inducing context while leaving model weights fixed, admitted changes remain inspectable, reversible, and compatible with closed-weight models.

Every Monday
Get next week’s papers.
Subscribe on Substack