Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

Tan Yu, Alexander Bukharin, Bryan Catanzaro, Jonathan Cohen, Jiantao Jiao and colleagues at NVIDIA (with the University of Minnesota and UC Berkeley) propose ways to predict which base checkpoint will become the strongest coding agent after agentic post-training.
Ask this paper
Problem. End-to-end pass@K on base models fails for agentic coding because many base checkpoints cannot reliably emit well-formed tool calls, while single-shot tasks skip the multi-step state tracking that matters.
Decisive step. Successful post-trained trajectories are replayed with tests rerun after each code edit to find the first step whose cumulative patch turns the repository from failing to passing.
Three screens. Decisive-Action bits per byte on the certified action, a Patch MCQ against verifier-rejected alternatives, and prefix-conditioned pass@K from the decisive context. None requires the base model to drive the harness from a cold start.
Results. Across ten public base and post-trained model pairs, all three screens rank the cohort in close agreement with post-trained SWE-bench Verified pass@1, while end-to-end base evaluation does not: only one base model resolves any task.
Reuse. The method needs only a benchmark's successful trajectories and its verifier, so future agentic coding benchmarks can be turned into base-model evaluations.
Abstract
How can we predict which base checkpoint is worth an expensive round of agentic post-training? End-to-end pass@$K$ tests whether successful behavior already appears in a base model's distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed tool invocation required to complete a task end-to-end. Single-shot or short-horizon tasks avoid these tool-calling failures by collapsing a multi-step interaction into a fixed prompt and a single patch, but they sidestep the core capability we care about: maintaining coherent state over many tool-using steps as the repository evolves. To bridge this gap, we treat successful post-trained agent trajectories as a lookahead signal of base-model potential. Replaying each trajectory and rerunning tests after every code-changing step identifies the decisive step: the first step whose cumulative patch flips the repository from failing to passing, certifying that the recorded action solves the task given the prior context. Motivated by a coverage principle for agentic traces, we build three screens at this step that do not require a base checkpoint to drive the harness from a cold start: (i) Decisive-Action BPB (bits per byte) measures the probability mass on the certified action, (ii) Patch MCQ tests the checkpoint's choice between that action and alternatives rejected by the same verifier, and (iii) prefix-conditioned pass@$K$ evaluates support for functionally-correct generations and credits any continuation that the tests accept. Across ten pairs of public base and post-trained models, all three screens rank the cohort in close agreement with post-trained SWE-bench Verified pass@$1$. As our methods need only a benchmark's successful trajectories and its verifier, they can be applied to turn future agentic coding benchmarks into base-model evaluations.