VERA: Scaling Verifiable Environments for Agentic co-Evolution

Junqi Liu, Yucheng Tang, Daguang Xu and colleagues at NVIDIA, with UC Santa Cruz, UIUC, NUS and Tsinghua, introduce VERA, which turns benchmark trajectories into 9,000+ verifiable sandboxes and uses them to co-evolve a model and its agent harness.
Ask this paper
Environment scaling. Benchmark trajectories become restartable, rubric-scored sandboxes for long-horizon tasks; an admission loop keeps only environments that execute and can be scored from observable evidence. The 9,000+ environment corpus is open-sourced.
Co-evolution. Both model weights and the harness are updated. A harness edit must pass self-tests and raise the development score by at least 5 points, and a model checkpoint is rejected if its score drops by more than 20%.
Results at 9B. The co-evolved agent beats the strongest single-axis baseline by 10.3 points on AutoCoWorkBench and 13.0 points on AutoMedBench.
Results at 27B. VERA reaches 71.6 on AutoCoWorkBench, above Claude Opus 4.8, and 80.7 on AutoMedBench, within 1.2 points of it; only at 27B are general capabilities retained and mostly improved.
Abstract
Competent agents need precise and verifiable environments, such as sandboxes that are resumable at any stage and evolve from observable evidence. However, most long-horizon work exposes how rare these are: for example, an agent in medical research must ground a finding, classify it, and write a report over dozens of dependent steps, yet recent environments score only the outcome. To address the challenges in stable training, we present VERA, which builds such environments at scale and lets agents evolve on them. VERA builds these environments from initial trajectories: an agent writes rubrics, executable checks, a judge verifies each sandbox, and only those that pass enter the training bank. On these environments, VERA alternates between two updates: train the model with rubric rewards, or edit the harness skills. We also create a verifier which gates model checkpoints and harness edits using explicit development-set acceptance criteria. This attribution distinguishes VERA's co-evolution from single-axis baselines: its updates target not only the cause but the outcome. With an open-source corpus of 9,000+ long-horizon verifiable environments, a 9B model paired with its co-evolved agent beats the strongest baseline by 10.3 and 13.0 points in the two domains. At 27B, it surpasses the baseline on AutoCoWorkBench (71.6) and AutoMedBench (80.7), transfers to unseen workflows, and retains general capabilities.