Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Minki Kang (KAIST, intern at NVIDIA), Byung-Kwan Lee, Yu-Chiang Frank Wang and colleagues at NVIDIA introduce Mid-Harness, which spends test-time compute at the boundary between model and harness: it samples several candidate shell actions and verifies them before one is executed.
Ask this paper
Motivation. In terminal tasks one bad command, such as a wrong package install, changes the environment and harms later steps even when the model could have produced a better command.
Verifier quality decides the gain. With a TMAX-9B generator, sampling more actions helps little under a weak verifier; a GPT-5.6 Sol verifier raises TerminalBench-Lite Pass@1 from 50.00% to 68.03% with 8 sampled actions.
Small-model verification. When TMAX-9B verifies its own candidates, pairwise comparison works best among the tested mechanisms, and distilling the stronger verifier into TMAX-9B improves Pass@1 further with the generator unchanged.
Cost. Combining action-level and trajectory-level scaling reaches higher success at lower estimated token cost than sampling more full trajectories, and the gains hold across additional models, benchmarks and harnesses.
Abstract
Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.