🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 26, 2026
Agents · Evaluation

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

First page
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
The curator’s take

Jingjie Ning, Xueqi Li and Yibo Kong of Carnegie Mellon, with Dongting Li of Tsinghua, introduce WhatWorkedBench, which scores research agents on whether they correctly predict how component changes affect results after a limited experiment budget.

Ask this paper

Key points
01

Task format. An agent inspects workflow code, chooses which configurations to run, and submits a response surface, a table predicting the score of every combination of component settings.

02

Ground truth. Exhaustive CPU execution gives the true effect of changing each component while holding the others fixed, across 36 tasks from 30 data sources, 8 workflow types and 1,248 configuration records.

03

Agent study. 108 agent episodes with DeepSeek V4 Flash and Pro plus 4,206 numerical-control records. With eight new measurements, a pair-effect ridge baseline picks the optimum on 15 of 22 sources.

04

Analysis beats raw agent guesses. Fitting a Gaussian process to the agent's own observations raises effect recovery from 0.632 to 0.698 in the original Flash cohort and from 0.621 to 0.720 in a second cohort.

05

Program structure helps. Encoding which configurations behave identically in code raises GP recovery from 0.248 to 0.462 on six workflows with six binary options.

Abstract

AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response surface, a table predicting scores for every configuration of component settings. Exhaustive CPU execution supplies reference effects for changing each component while holding the others fixed. These effects capture combinations of changes across 36 tasks from 30 data sources and 8 workflow types, with 1248 configuration records. Core evaluation combines 4,206 numerical-control records across all eight families and 108 agent episodes across the original six. At eight new measurements, pair-effect ridge selects an optimum on 15 of 22 sources and limits every effect error to 10% of score range on three. Fitting a Gaussian process (GP) to the same agent observations raises effect recovery, accuracy relative to true effect magnitude, from 0.632 to 0.698 in the original Flash cohort and from 0.621 to 0.720 in an additional cohort. On six completed beat-detection and graph submissions, the same-observation GP raises family-macro recovery from 0.303 to 0.455. On six workflows with six binary options at 20 new measurements, encoding code equivalences, configurations with identical behavior, raises GP recovery from 0.248 to 0.462. WhatWorkedBench supports research on experimental agents, adaptive experimental design, numerical inference, and use of program structure.

Every Monday
Get next week’s papers.
Subscribe on Substack