TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

Radhika Gaonkar (Prime Intellect) introduces TRACE, a protocol that tests whether a change in an agent's verifier score reflects a change in the agent or a change in the evaluation.
Ask this paper
Protocol. TRACE applies a targeted change to one part of an evaluation, compares paired runs, checks whether agent behavior changed, and rescores unchanged trajectories to isolate the scoring rule.
Synthetic suite. On 25 synthetic tasks, renaming tools lowers a scripted agent's score by 0.250 although it performs identical operations; restoring the original names at scoring time removes the entire gap.
tau2-bench. On 88 new tau2-bench tasks with four LLM agents and repeated runs, renaming tools or reformatting outputs moves reward by less than 0.10 for seven of eight agent-change pairs, while deliberately misleading tool names lower every agent's reward by 0.20 to 0.44.
Run-to-run noise. Identical reruns flip 15-36% of task outcomes, so single-run comparisons cannot separate presentation effects from variance. An initial 30-task study produced one clear effect that did not replicate.
Judges. Two frontier LLM judges are each stable when a fixed trajectory is presented differently, yet disagree with each other on 57% of records, mostly because one grades procedure rather than outcome.
Abstract
Verifier scores now serve as both benchmark metrics and training rewards for large language model (LLM) agents, and a change in score is routinely read as a change in capability. It may instead reflect a change in the evaluation. We introduce TRACE, a protocol that turns a score change from a verdict into a testable diagnosis: it applies a targeted change to one part of an evaluation, compares paired runs, checks whether the agent's behavior changed, and rescores unchanged trajectories to test whether the scoring rule is responsible. In a controlled suite of 25 synthetic tasks, renaming tools lowers a scripted agent's score by 0.250 even though it performs exactly the same operations; restoring the original names at scoring time closes the entire gap, while the same mutation exposes a genuine behavioral failure in a second agent. On public $\tau^2$-bench tasks with four LLM agents, an initial 30-task study finds mixed reward changes whose one clear effect does not replicate. In a larger follow-up on 88 new tasks with repeated runs per condition, renaming tools or reformatting tool outputs leaves reward unchanged to within $\pm$0.10 for seven of eight agent-change pairs, whereas tool names that deliberately mislead lower every agent's reward by 0.20-0.44, showing that the setup can detect real effects. Identical reruns flip 15-36% of task outcomes, so single-run comparisons cannot separate presentation effects from run-to-run variation. Two frontier LLM judges give consistent verdicts when a fixed trajectory is presented differently, yet disagree with each other on 57% of the same records, largely because one grades procedure rather than outcome. TRACE thus separates what a score change says about the agent from what it says about the measurement.