🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Efficiency · Evaluation

The Harness Effect

First page
The Harness Effect
Paper summary

As orchestration harnesses mediate every model call, this study asks how much the harness alone moves cost and performance. It ran 22 evaluation tasks across six foundation models, then changed only the orchestration layer while holding the models constant.

Ask this paper

Key points
01

Harness-only, models fixed: By varying just the orchestration layer over models like Claude Sonnet 4.6, Gemini 3.1, Qwen 3.6, and GLM 5.1, the study isolates the harness as the variable and measures its independent effect.

02

Big, consistent savings: Holding models constant, the harness cuts blended cost per task 41%, tokens per task 38%, and median wall-clock 44%, with completion quality at parity.

03

Two clean regularities: Efficiency is model-invariant, every model gets 33 to 61% cheaper, while quality gain correlates almost perfectly with baseline model strength (r=0.99), an effect the authors call harness leverage.

04

Why it matters: On this workload the orchestration layer moved cost per task more than the entire spread of the model menu did, making the harness the one component whose efficiency multiplies across every model a team runs.

Every Monday
Get next week’s papers.
Subscribe on Substack