🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 6, 2026
Agents

When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge

First page
When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge
The curator’s take

Xi Qin, Isabel Kurth and colleagues at SAP Lab build a meta-agent pipeline in which Claude Opus generates terminal tasks and verifiers for RL, and document why RL training of a 9B model still stalls.

Ask this paper

Key points
01

Dataset. 18,185 terminal tasks in the Harbor format, generated by Claude Opus 4.6 and 4.8 with prompts designed around casual engineer voice, anti-leakage rules that forbid restating verification criteria, and domain coverage.

02

Three failure classes. The authors separate benchmark invalidity, harness brittleness and reward misalignment, and show that a runnable Docker image and passing test suite do not guarantee a faithful training pipeline.

03

Solvability band. Prompt redesign and longer context raise baseline solvability 5.6x, but Qwen3.5-9B then saturates at 81.3% mean pass@2 within 20 steps; adding hard tasks drops it to 20.6% with the same configuration. At about 10% solvability, most batches contain no success and validation pass@1 stays between 3.9% and 5.9%.

04

Takeaway. The usable range of task difficulty depends on the model being trained, so task generation has to be calibrated to the student rather than to the generator.

Abstract

Using a frontier model like Claude Opus as a meta-agent to generate terminal tasks and verifiers for RL training is increasingly common. Yet a runnable Docker image and executable test suite do not guarantee a faithful end-to-end pipeline for terminal agent training. We present a meta-agent pipeline motivated by this gap, diagnosing three classes of failure: benchmark invalidity, harness brittleness, and reward misalignment. Prompt redesign and context extension raise baseline solvability 5.6 times, but a 9B model saturates at 81.3% mean pass@2 within 20 steps on Claude Opus-generated tasks. Adding hard tasks reduces mean pass@2 to 20.6% without changing the training configuration, a strong evidence that the solvability band is model-specific. These findings demonstrate that meta-agent reliability requires solvability-band calibration, verifier audits, and infrastructure error accounting as first-class evaluation criteria, not post-hoc diagnost.

Every Monday
Get next week’s papers.
Subscribe on Substack