Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale

Hao Fu, Baiting Zhu, Minglei Chen, Yinjie Huang and Shuai Ding (Meta) describe EvoPilot, a human-gated method for running LLM-agent research loops on a production retrieval system, and report a 37-day campaign on the retrieval stack behind Video Deep Dive.
Ask this paper
Why online autoresearch is different. Production experiments are asynchronous, take hours per variant and run for weeks, so a finished run can still support a wrong conclusion through a no-op code change, leaked data windows, drifting evaluator semantics or arms that traverse different serving funnels.
Method. Role-specific agents run each round through a versioned domain skill and a typed adapter, durable records keep experiments and failures, and deterministic checks enforce lessons already recorded.
Caught a false negative. A simple autoresearch loop blamed a 22-point offline hit-rate drop on an interaction head; EvoPilot's verification traced it to a pre-existing evaluation defect, and after the fix the head measured a 3.20-point offline gain.
Online result. A seven-day randomized online test estimated a 0.66% relative increase in the Video Deep Dive slice of Good Search Result Rate for Retention.
Operational checks. Replay and mutation tests rejected invalid comparisons while admitting valid ones, durable state recovered an interrupted round, and artifact reuse saved about five GPU-hours.
Abstract
Large language model (LLM) agents can propose, implement, and evaluate model changes. Autoresearch loops demonstrate this capability through minutes-scale iterations on a self-contained program. Online autoresearch instead spans asynchronous systems, hours-long variants, and weeks-long campaigns that can influence a product. A completed run can still support an invalid conclusion when a code change is a no-op, data windows leak, evaluator semantics drift, or the two arms traverse different serving funnels. We present EvoPilot, a human-gated method for long-horizon online autoresearch. Role-specific agents execute each round through a versioned domain skill and typed adapter. Durable records preserve experiments and failures; deterministic checks enforce recorded lessons. We study a 37-day campaign for the retrieval system that powers Video Deep Dive (VDD), an online experience for discovering follow-on videos after a user opens a seed video. The campaign covered seven directions and used an hourly refreshed index of hundreds of millions of videos. Earlier manual experiments had not established a benefit from an interaction head. A primitive autoresearch attempt revisited the direction but incorrectly attributed an offline hit-rate decline of 22 percentage points to the head. We then introduced EvoPilot. Its human-gated verification traced the drop to a pre-existing evaluation defect that produced output depths of 3,000 and 600. After repair, a matched comparison measured an offline improvement of 3.20 percentage points. Post-study replay and mutation tests rejected invalid comparisons while admitting valid counterparts. Durable state recovered an interrupted round, and artifact reuse avoided approximately five GPU-hours. Separately, a seven-day randomized online evaluation estimated a 0.66% relative increase in the VDD slice of Good Search Result Rate for Retention (GSRR).