PatchBench: Evaluating AI Agents for Vulnerability Patching

Chihao Shen and colleagues at Maryland and UC Davis show that PoC-only validation inflates vulnerability-patching solve rates by 1.83x on average, because agents either recall the historical developer patch or fix the crash rather than the bug.
Ask this paper
Two distinct validity threats, both measured: a patch similarity metric finds that 25% of agent patches are substantially similar to historical developer patches, and agents frequently patch on the crash stack trace to suppress the symptom.
PatchBench is built to defeat both: it selects vulnerabilities whose ground-truth fixes lie outside the crash stack, then uses vulnerability transplant and code mutation to move historical CVEs into new repository contexts.
New validation checks security and semantics: the benchmark scores whether the patch is both actually secure and semantically correct, not merely non-crashing.
11 state-of-the-art agents including the top three AIxCC entrants, so the inflation finding covers the systems people cite as the state of the art.
Why it matters: memorization contaminating security benchmarks is the same disease as data contamination in coding evals, but the downstream cost of a false pass is much higher.
Abstract
AI agents have recently demonstrated strong performance in automated vulnerability patching. However, existing evaluations often validate a patch only by testing whether the provided Proof-of-Concept (PoC) input still triggers a crash. This leaves two key threats to validity: agents may reproduce memorized historical developer patches, or they may generate surface-level fixes that only suppress the reported crash. We study these concerns for C/C++ vulnerability patching. We introduce a patch similarity metric to detect memorized patches. On average, 25% of the agent patches exhibit substantial similarity to historical developer patches, indicating that patch memorization is a real threat to the validity of vulnerability patching evaluations. Meanwhile, agents also frequently exploit benchmark structures to pass patch validation by patching on the crash stack trace to suppress the crash, rather than localizing and fixing the root cause of the vulnerabilities. To handle these issues, we propose PatchBench, a new benchmark for evaluating AI agents on realistic vulnerability patching tasks. PatchBench selects vulnerabilities whose ground-truth fixes lie outside the crash stack and uses vulnerability transplant and code mutations to migrate historical vulnerabilities into new repository contexts, reducing the risks of surface-level fixes and patch memorization. We develop new patch validation methods that thoroughly evaluate both security and semantic correctness of agent patches. Across 11 state-of-the-art agents, including the top three AIxCC agents, the original PoC-only validation inflates the patching task solve rate of agents by 1.83$\times$ on average. Our results reveal key limitations of current patching agents and point to future research directions for more reliable vulnerability repair.