🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Reasoning · Agents

Verification as a Scaling Axis

First page
Verification as a Scaling Axis
Paper summary

Verification is emerging as a distinct scaling axis alongside pre-training and test-time compute, and this Stanford, NVIDIA, and UC Berkeley collaboration builds a training-free verifier that reads a continuous, calibrated score straight off the scoring-token logits instead of trusting a discrete pass or fail grade.

Ask this paper

Key points
01

Scores from logits, not grades: Rather than asking a judge model for a discrete verdict, the method reads a continuous calibrated score off the scoring-token logits, giving a smoother and more informative signal with no fine-tuning.

02

Three tuning knobs: Accuracy improves through score granularity for cleaner separation, repeated evaluation to cut variance, and criteria decomposition to reduce complexity, all without touching model weights.

03

Broad, strong numbers: It reaches 86.5% on Terminal-Bench V2, 78.2% on SWE-Bench Verified, 87.4% on RoboRewardBench, and 73.3% on MedAgentBench, spanning coding, robotics, and medical agents.

04

Why it matters: The same continuous score doubles as a dense reward for SAC and GRPO and as a task-progress signal shipped in a Claude Code extension, so one verifier serves evaluation, training, and live agent monitoring at once.

Every Monday
Get next week’s papers.
Subscribe on Substack