How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

Kirill Brilliantov, Alejandro Hernandez-Cano and Emmanuel Abbe at EPFL and Apple test whether elaborate MLE-agent harnesses help strong models, comparing them under matched budgets with Malena, a single-session coding agent that has only read, write and bash tools.
Ask this paper
Controlled comparison. Same frontier backbone and the same time budget for all systems, addressing MLE-bench results that are usually reported with different backbones, hardware and few seeds.
Main result. Open-source state-of-the-art harnesses (multi-agent orchestrators, retrieval subagents) show no advantage over the minimal agent; Malena matches or beats four of them on MLE-bench and NatureBench at every frontier backbone tested. The one significant regression is on the weakest backbone, Gemma 4 31B, where MLEvolve scores higher on the percentile metric.
Trace analysis. The minimal agent performs on its own the search those harnesses hard-code, balancing rare techniques against refinement and reusing artifacts it already produced.
Conclusion. For current MLE benchmarks, performance is driven mainly by the backbone, and effort spent on hand-built harness layers gives little return.
Abstract
Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orchestrators, dedicated retrieval subagents, and more. While such harnesses expand, the use of more primitive but improved coding agents - where LLMs have direct access to the execution environment through read, write, and bash primitives - has received little attention in the field. In this paper we find that, under an equal time budget and the same frontier LLM backbone, open-source state-of-the-art harnesses provide no advantages over a single session of a minimal-harness coding agent baseline, pointing to the backbone as the primary driver for performance. Via a series of large-scale systematic ablation studies, we argue that the machinery layers become redundant in the coding agent setting. We conclude that the effort spent elaborating hand-crafted harnesses around strong models yields poor returns for current MLE benchmarks.