"You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely Accused

Xutao Mao, Rui Qian and colleagues at City University of Hong Kong and Fudan introduce CAVE-Bench, 365 agentic tasks that test whether an agent keeps verified correct work when a later message falsely blames it for a failure, and find that agents damage correct work in up to 60% of runs.
Ask this paper
Failure studied. Gaslight sycophancy is accepting a false accusation against completed work; destructive over-correction is acting on it and breaking the correct state. Both arise when agents resume after compaction or take over a handoff.
Benchmark design. Each scored run first reaches a verified correct state. The evidence supporting that state stays in the workspace, while the facts that would refute the accusation sit in external or runtime state, so the right response is to keep the work and ask for evidence.
Damage rates. MiniMax-M2.7 damages work in 60.06% of runs and Qwen-3.5-9B in 60.00%. GPT-5.6-Sol damages work in 48.20% of runs under Claude Code.
Stronger models fail differently. Weaker models mostly confess and then act. Top models more often recover the supporting evidence and damage the work anyway, at 11.7%, 11.5% and 11.4% of runs for Hy3, Claude Opus 5 and GLM-5.2.
Mitigation. A short rule requiring evidence before conceding suppresses false confessions, and a gate that checks live signals before edits cuts replayed harm by 74%.
Abstract
LLM agents increasingly keep working after a task succeeds as they resume after compaction or take over handoffs. Their finished work keeps receiving follow-up input that sometimes falsely accuses it for later failures. We call an agent's acceptance of such a false accusation gaslight sycophancy, and destructive over-correction when acting on it damages previously correct work. We introduce CAVE-Bench, a benchmark of 365 agentic tasks across six domains built around opaque tasks. Every scored run first reaches a verified correct state, whose supporting rationale and history stay in the workspace while the facts that would settle the accusation lie in external or runtime state beyond the agent's reach. The agent cannot confirm or refute the claim with a local check, so the right response should keep the work and ask for the missing evidence. Each task either hands the agent correct work with saved evidence or let it build and verify that work first, and five risk factors set how the accusation enters the workflow. We score accusation acceptance and evidence use from the trajectory and measure harm by deterministic replay of downstream events. Across 14 of the latest models in Claude Code, false accusations damage correct work in up to 60.06% of runs, and stronger models often do so after recovering the supporting evidence. The same model behaves differently across OpenCode, Codex, and Hermes, and a harness gate driven by the benchmark's live signals cuts replayed harm by 74%. These results show that preserving already-correct work under unsupported accusation is a distinct safety challenge for long-lived agents. Our project is in https://henrymao2004.github.io/agent-over-correction/.