Verification as a Scaling Axis

Verification is emerging as a distinct scaling axis alongside pre-training and test-time compute, and this Stanford, NVIDIA, and UC Berkeley collaboration builds a training-free verifier that reads a continuous, calibrated score straight off the scoring-token logits instead of trusting a discrete pass or fail grade.
Ask this paper
Scores from logits, not grades: Rather than asking a judge model for a discrete verdict, the method reads a continuous calibrated score off the scoring-token logits, giving a smoother and more informative signal with no fine-tuning.
Three tuning knobs: Accuracy improves through score granularity for cleaner separation, repeated evaluation to cut variance, and criteria decomposition to reduce complexity, all without touching model weights.
Broad, strong numbers: It reaches 86.5% on Terminal-Bench V2, 78.2% on SWE-Bench Verified, 87.4% on RoboRewardBench, and 73.3% on MedAgentBench, spanning coding, robotics, and medical agents.
Why it matters: The same continuous score doubles as a dense reward for SAC and GRPO and as a task-progress signal shipped in a Claude Code extension, so one verifier serves evaluation, training, and live agent monitoring at once.