RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

Zheng Chen, Linfeng Liu, Hong Li and Hong Yan at Meta present RankEvolve, an auto-research framework for generative ranking models that treats execution accuracy (whether each code change is implemented correctly) as the main constraint on long research runs.
Ask this paper
Failure model. A single silent defect, such as leaked held-out data, a missing layer norm, a disconnected gradient or an unwired train/eval flag, wastes hours of accelerator time and produces an invalid metric that later iterations build on.
Executable Operating Protocol. Phases, gates, branches and loops are declared in an EOP and compiled to a state machine that the runtime enforces, instead of being left to the agent's prompt.
Composing coding-agent products. A meta-meta-harness uses complete products such as Claude Code and Codex as graph nodes that review and repair one another; in a budget-matched test this raises all-oracle execution accuracy from 45.8% (best single product) to 62.5% (+16.7 points, 95% CI 6.6 to 26.7) with a 10.4% silent critical-defect rate.
Deployment. Twelve iterations on the open-source HSTU recommender reached NDCG@10 of 0.2192 on MovieLens-20M LARGE (+4.48% over the published anchor); a LitGPT split reproduces the composition effect (+12.5 points).
Abstract
Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leave a train/eval flag unwired, invalidating expensive runs and compounding error across iterations. We present RankEvolve, an auto-research framework for evolving generative ranking models. An Executable Operating Protocol (EOP) declares phases, gates, branches, and loops, and the runtime enforces the compiled state machine. A meta-meta-harness composes complete black-box coding-agent products, including Claude Code and Codex, as execution-graph nodes that review and repair one another's work. In a budget-matched evaluation, heterogeneous composition raises all-oracle execution accuracy from the best single-product baseline of 45.8 percent to 62.5 percent (paired +16.7 points, 95 percent CI [6.6, 26.7]) while achieving a 10.4 percent silent critical-defect rate. An implemented knowledge layer carries findings, including negative results, across iterations. In a twelve-iteration deployment on the open-source HSTU recommender, RankEvolve reported NDCG@10 of 0.2192 on MovieLens-20M LARGE (+4.48 percent over the published anchor) and 0.1948 on BASE (+2.80 percent). ExecML-HSTU, seeded by incidents from that deployment, provides the oracle benchmark for the execution-accuracy evaluation. A pre-specified LitGPT transfer split replicates the heterogeneous-composition effect beyond recommendation (+12.5 points, 95 percent CI [3.0, 22.0]), and a paired ablation isolates per-step from full-protocol instruction injection. These results characterize when runtime-controlled composition of coding-agent products improves execution accuracy.