🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 1, 2026
Agents · Memory

GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements

First page
GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements
The curator’s take

Zhibang Yang and colleagues at Peking University introduce GitHarness, which stores an agent's requirement states and work states as a branchable Git-style history so that when a user adds, changes or revises a requirement, the agent restores the right earlier state instead of rewriting everything.

Ask this paper

Key points
01

Problem. Mid-task requirement changes usually touch part of the accumulated work, yet agents either keep obsolete material or redo everything.

02

Git Agent. A trainable agent resolves each requirement change, picks a compatible historical state, and the version interface restores it on a new branch so the underlying harness inherits valid work and re-executes only affected parts. It is trained with interface-level black-box RL while the harness and executor stay fixed.

03

MTAgentBench. A new benchmark with requirement-change trajectories across math, text-to-SQL, agentic search, software engineering and research synthesis.

04

Results. On top of TCRAG, StackPlanner and OpenHands, GitHarness gives the best scores in most cells; with Qwen3-32B on TCRAG, math rises from 72 to 89 and research synthesis from 25.0 to 38.5, and with DeepSeek-V4-Flash agentic search rises from 34 to 56. The RL-trained variant does not beat the untrained one in most cells.

Abstract

LLM-based agents increasingly collaborate with users on long-horizon tasks, accumulating evidence, code, and drafts through extensive search, reasoning, and execution. As users inspect these results, they may supply missing information requirement completion, introduce new requirements requirement elicitation, or revise existing ones requirement shift. These changes often affect only part of the accumulated work, yet agents may carry forward obsolete information or turn local revisions into global rewrites. Existing approaches clarify current intent without determining how prior work should change, or reuse execution histories under a fixed objective. We address this gap by formulating dynamic-requirement collaboration as joint requirement tracking and local update. We introduce GitHarness, a pluggable Git-style framework that organizes requirement states and their corresponding harness work states into a branchable version history. A trainable Git Agent resolves requirement changes and selects a semantically compatible historical state. A unified version interface then restores that state and creates a new branch, enabling the underlying harness to exclude obsolete information, inherit compatible work, and focus execution on affected parts. The Git Agent is trained through interface-level black-box reinforcement learning, with downstream harnesses and task-execution models kept fixed. We also construct MTAgentBench, a verifier-preserving benchmark covering mathematical reasoning, text-to-SQL, agentic search, software engineering, and research synthesis. Experiments demonstrate strong task performance alongside effective requirement tracking, preservation of valid work, and efficient execution.

Every Monday
Get next week’s papers.
Subscribe on Substack