🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 8, 2026
Robotics · Reasoning

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

First page
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
The curator’s take

Zewei Zhou, Boris Ivanovic, Marco Pavone and colleagues at NVIDIA with UCLA, UC Berkeley and Stanford introduce VeriFine, a harness that improves the policy, its training curriculum and its judge together for embodied reasoning.

Ask this paper

Key points
01

Policy loop. A rubric judge diagnoses recurring failures, builds an adaptive curriculum and drives policy optimization.

02

Judge loop. When progress stalls, the system asks humans for guidance on informative failure cases, and humans and agents resolve disagreements to recalibrate the judge toward a physical-reasoning rubric.

03

Results. On driving, the policy reasoning score rises from 60.56 to 71.61 over three rounds, an 18.2% relative gain under a fixed reference evaluator; judge-guided data selection beats random selection by 4.5% to 14.6% per round.

04

Judge quality. The final judge reaches about 0.85 correlation with humans, and gains hold for both RL and supervised fine-tuning and for robot navigation.

Abstract

Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.

Every Monday
Get next week’s papers.
Subscribe on Substack