The Harness Effect
Free while signed in. Answers cite the passages they came from.

As orchestration harnesses mediate every model call, this study asks how much the harness alone moves cost and performance. It ran 22 evaluation tasks across six foundation models, then changed only the orchestration layer while holding the models constant.
Harness-only, models fixed: By varying just the orchestration layer over models like Claude Sonnet 4.6, Gemini 3.1, Qwen 3.6, and GLM 5.1, the study isolates the harness as the variable and measures its independent effect.
Big, consistent savings: Holding models constant, the harness cuts blended cost per task 41%, tokens per task 38%, and median wall-clock 44%, with completion quality at parity.
Two clean regularities: Efficiency is model-invariant, every model gets 33 to 61% cheaper, while quality gain correlates almost perfectly with baseline model strength (r=0.99), an effect the authors call harness leverage.
Why it matters: On this workload the orchestration layer moved cost per task more than the entire spread of the model menu did, making the harness the one component whose efficiency multiplies across every model a team runs.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack