Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

Xing Zhang, Peiyang He and colleagues at AWS Forward Deployed Engineering make the verifier itself the evolving object in a self-improving agent loop, building inspectable graders from small deterministic drawback detectors. Accepted at the NeurIPS 2026 workshop Who Verifies the Agents?
Ask this paper
Verifier design. The verifier is an inspectable expression over mostly deterministic drawback detectors synthesized from clustered failures, gated at creation and selected for agreement with a ten-item anchored reference set plus consensus on unlabeled outputs, never for the agent's score.
Agreement gain. On MBPP+ the evolved verifier gains +0.21 held-out agreement over the hand-authored seed composition on every seed and ends ahead of the bare LLM judge it contains.
Key finding. Removing the anchor guards collapses the verifier into an always-pass grader, and that collapsed verifier trains skills just as well, so downstream task score cannot certify a self-evolved verifier.
Double Ratchet. Pairing the evolved verifier with a lifecycle-managed skill loop retains 88-110% of the lift that ground truth or a rubric gives the same loop, across code generation, enterprise text-to-SQL and reference-free report generation.
Rubric gaming. When evolved skills gamed the report rubric, an outer judge caught it and one added detector repaired it, but the judge was itself wrong until it was given the task contract.
Abstract
We changed the agent: did it actually get better? Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier. On open-ended tasks none exists, so the loop is handed a hand-written rubric or a bare LLM judge grading output from a model like itself, inviting reward hacking and shared blind spots. We make the verifier the evolving object: an inspectable expression over small, mostly deterministic drawback detectors, synthesized from clustered failures, gated at birth, and selected for agreement with a ten-item anchored reference set plus consensus over unlabeled outputs, never for the agent's score. On MBPP+ it gains +0.21 held-out agreement over the hand-authored seed composition, on every seed, and ends ahead of the bare LLM judge it contains. One finding should change how co-evolved verifiers are validated: removing the anchor guards collapses the verifier into a vacuous always-pass grader, yet that collapsed verifier trains skills just as well. Downstream task score cannot certify a self-evolved verifier. Score does answer sufficiency, and there an evolved verifier can substitute: Double Ratchet, pairing the verifier with a lifecycle-managed skill loop, retains 88-110% of the lift that ground truth or a rubric buys the same loop, across code generation, enterprise text-to-SQL, and reference-free report generation. When evolved skills gamed the report rubric, an outer judge caught it and one added detector repaired it; the judge itself was wrong until given the task contract.