🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 1 – Sep 1, 2026
Agents · Code

Can escalation channels redirect reward hacking toward defect disclosure?

First page
Can escalation channels redirect reward hacking toward defect disclosure?
The curator’s take

Francesca Gomez gives coding agents a structured way to report broken test infrastructure at the moment of conflict, and reward hacking drops from 23.6% to 5.3% with no measured performance cost.

Ask this paper

Key points
01

Redirect capability rather than contain it: The same ability that lets an agent detect and exploit a defect lets it report one. The intervention changes the decision environment instead of restricting the agent.

02

Clean 2x2 design: Escalation tool, standalone anti-reward-hacking policy, and their combination, separated factorially across 8 frontier models from 5 families.

03

The effect size: Combined intervention takes reward hacking from 23.6% to 5.3%, mixed-effects logistic OR 9.2 (95% CI 5.0 to 16.8, p < 1e-12), eliminating it entirely for 6 of 8 models with no detectable cost or overhead.

04

Escalation and hacking are near mutually exclusive: 96.8% of escalations involve no hacking, which is what makes escalation a usable signal rather than a second channel to game.

05

Diagnostic value on top: Escalation adds 10.1 percentage points of defect detection coverage beyond monitoring and is more accurate once it fires, 99.4% against 85.8%. Read alongside BAITBENCH, where a policy line alone barely moved the rate.

Abstract

When coding agents encounter defective test infrastructure they may reward-hack: hardcoding outputs or editing test files to pass tests they cannot legitimately satisfy, a pattern that has now appeared outside benchmarks, in a coordinated multi-agent intrusion of a major AI platform's production infrastructure. The same capability that lets an agent detect and exploit a defect could let it report one, given the right decision environment. We evaluate escalation channels, structured reporting tools available to the agent at the point of conflict, as a decision-environment intervention that both reduces reward hacking and surfaces the infrastructure defects that trigger it. A $2 \times 2$ factorial separates the contributions of an escalation tool, a standalone anti-reward-hacking policy, and their combination. Across 8 frontier models spanning 5 families, the combined intervention reduces reward hacking from 23.6\% to 5.3\% (mixed-effects logistic OR = 9.2, 95\% CI 5.0--16.8, $p < 10^{-12}$) with no detectable cost or performance overhead, eliminating it entirely for 6 of 8 models. Escalation and hacking are near-perfectly mutually exclusive, with 96.8\% of escalations involving no hacking. Beyond reduction, escalation channels function as diagnostic infrastructure: on top of monitoring, escalation adds +10.1 percentage points of defect detection coverage and is more accurate once it fires (99.4\% vs 85.8\%). Unlike containment-based approaches that risk outpacing growing model capabilities, escalation channels redirect capability toward disclosure rather than exploitation.

Every Monday
Get next week’s papers.
Subscribe on Substack