🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 5, 2026
Agents

VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses

First page
VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses
The curator’s take

Jiexing Qi and colleagues at Huawei's ICT AI Competence Center propose VACE, which alternates agentic RL on the model with trajectory-driven revisions to the harness, and accepts a harness revision only if it improves validation performance with the newly trained model.

Ask this paper

Key points
01

Loop. After each RL stage, the collected trajectories are used to propose a harness revision. The incumbent and candidate harness are both evaluated with the updated model held fixed, and the candidate guides the next round of training only if it wins on validation.

02

Results with Qwen3.5-9B. 45.26% test accuracy on OfficeQA and 75.19% mean partial credit on AutomationBench.

03

Against ablations. That is 6.43 and 9.09 points above weight-only RL, and 4.59 and 6.95 points above alternating without the validation gate.

04

Why the gate matters. 17 of 44 harness proposals lowered validation performance at the current checkpoint and were rejected before they could shape training.

Abstract

Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training. We propose VACE, Validation-Gated Alternating CoEvolution, which alternates agentic reinforcement learning with trajectory-driven harness refinement. After each RL stage, VACE reuses the collected trajectories to propose a harness revision and evaluates the incumbent and candidate with the updated model held fixed. The candidate guides subsequent training only if it improves validation performance. With Qwen3.5-9B, VACE achieves 45.26% test accuracy on OfficeQA and a mean partial-credit score of 75.19% on AutomationBench, exceeding weight-only RL by 6.43 and 9.09 percentage points and ungated alternation by 4.59 and 6.95 points, respectively. Across 44 harness proposals, 17 reduce validation performance at the updated checkpoint and are rejected before subsequent RL training, highlighting the importance of validation gating.

Every Monday
Get next week’s papers.
Subscribe on Substack