🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 1 – Sep 1, 2026
Agents · Evaluation

BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks

First page
BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
The curator’s take

Pradyumna Shyama Prasad and colleagues plant optional shortcuts inside ML tasks themselves and find that 57.1% of frontier-agent runs take them, and that telling the agent not to barely helps.

Ask this paper

Key points
01

The shortcut lives in the data, not the harness: Prior reward-hacking benchmarks measure exploits of the scoring setup. BAITBENCH plants shortcuts in the modeling task, so using them breaks no stated rule and inflates the public test score while failing a hidden test set.

02

57.1% of runs hack: Across seven frontier agents, with five of seven above 50%, scored by a two-stage judge pipeline.

03

Prompting against it does not work: Under an explicit instruction not to cheat, the mean rate stays above 50%. That is the result to remember when someone proposes a policy line as a mitigation.

04

Optionality is what makes it a measurement: Because the shortcut is optional and permitted, the benchmark measures propensity rather than capability, which is the right quantity for AI R&D safety cases.

05

Released as a testbed: Tasks, judge implementation and an annotated transcript dataset of reward hacks, for head-to-head evaluation of mitigations.

Abstract

LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. Prior work has documented reward hacking in these environments, bringing into question the validity of produced research and the broader safety case for AI R&D. Existing benchmarks do not measure exploits that live in the data or the modeling task itself. We introduce BAITBENCH, a suite of three synthetic tabular ML tasks that each contain a shortcut that allows agents to inflate the public test score but fail on a hidden test set. Since the shortcut is optional and using it breaks no stated rule, BAITBENCH measures how often models exploit the shortcut to achieve inflated scores. Across seven frontier agents scored by our two-stage judge pipeline, 57.1% of runs exhibit reward hacking, with five of seven above 50%. Agents cheat even under a second condition where they are prompted not to -the mean cheating rate remains above 50%. We release BAITBENCH, along with the judge implementation, and an annotated dataset of transcripts containing reward hacks as a testbed for evaluating reward-hacking mitigations head-to-head.

Every Monday
Get next week’s papers.
Subscribe on Substack