Finding the Right Fit: Model-Harness Interactions across Agent Tasks

Yixuan Li, Bo An and colleagues at Nanyang Technological University evaluate 66 model-harness configurations and show that model rankings, best harnesses and cost-efficiency all change with the harness and the benchmark, so the pairing has to be evaluated as a unit.
Ask this paper
Setup. Four configurable harnesses (OpenHands, DeepSeek Harness, PI, openJiuwen) paired with five models on TUA-Bench, ALE-CLI and Terminal-Bench 4, plus the native Codex-GPT and Claude Code-Claude pairings.
Rankings reverse. On Terminal-Bench 4, Claude leads GPT by 7.94 points in OpenHands and trails it by 30.16 points in PI.
No universal best harness. For four of five models the best harness changes between benchmarks. One stable pairing exists: openJiuwen gives Kimi its highest score on all three benchmarks, by 5.61 to 11.11 points.
Vendor harness and cost. A model's own vendor harness is not reliably its best, and spending more does not reliably score higher; GPT on Terminal-Bench 4 scores higher under PI than under DeepSeek Harness at under a quarter of the cost per task.
Why fit varies. Models initiate almost all repairs themselves, so the deciding factor is whether the harness returns failures in a form the model can use. GPT does best in PI's minimal scaffold; Kimi, which often emits malformed tool calls, does best in openJiuwen. All 6,204 scored trajectories are released.
Abstract
Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when the setting changes. We evaluate 66 configurations: four configurable harnesses (OpenHands, DeepSeek Harness, PI, and openJiuwen) paired with five models on TUA-Bench, ALE-CLI, and Terminal-Bench 4, plus the native Codex-GPT and Claude Code-Claude pairings. Model rankings reverse across harnesses. On Terminal-Bench 4, Claude leads GPT by 7.94 points in OpenHands but trails it by 30.16 points in PI. For four of the five models, the best harness changes from one benchmark to another, yet some pairings hold: openJiuwen gives Kimi its highest score on all three benchmarks, by 5.61 to 11.11 points. A model's own vendor harness is not reliably its best, and higher cost does not reliably buy a higher score. On Terminal-Bench 4, GPT scores higher under PI than under DSH at less than a quarter of the cost per task. Matched trajectories suggest why fit varies. Models start almost all repairs themselves, so much depends on whether the harness hands failures back in a form the model can use. GPT does best with PI's lean scaffold, while Kimi, which often issues malformed tool calls, does best in openJiuwen. We argue that the model, the harness, and the task should be evaluated together, and we release the harness adapters, evaluation code, and all 6,204 scored trajectories at https://github.com/liyix/finding-the-right-fit and https://huggingface.co/datasets/yixuanli97/finding-the-right-fit.