🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 9, 2026
Reasoning

Finding Blind Spots in AppWorld and WorkArena Task Verifiers

First page
Finding Blind Spots in AppWorld and WorkArena Task Verifiers
The curator’s take

Richard Abrich (OpenAdapt.AI) audits the shipped task verifiers of AppWorld and WorkArena for false accepts using source-informed mutation tests. Accepted at a NeurIPS 2026 workshop on verifying agents.

Ask this paper

Key points
01

Gap. Published audits mostly sample FAIL verdicts to estimate false rejects. This one constructs wrong task effects and checks whether verifiers still return PASS.

02

AppWorld. Duplicating a non-idempotent write creates an extra record while keeping every checked field correct. The verifier accepts all three task variants from two of five eligible generators (6 of 15 constructed effects). A cardinality check added to copies of the checker makes all six fail.

03

WorkArena. Of 23 extra-field cases rerun on a hosted ServiceNow instance, independent Table API readback confirms non-default persisted values in 21, and all 23 receive PASS.

04

Scope. These are selected cases that confirm wrong effects; they do not estimate a population rate. Intent-swap grids produced no PASS across 2,689 off-diagonal executions.

05

Why it matters. Leaderboards and verifier-reward RL loops inherit every false accept.

Abstract

Execution-based task verifiers decide whether an agent succeeded. We audit shipped AppWorld and WorkArena verifiers with source-informed mutation tests. The main audit never modifies a shipped checker. In AppWorld, duplicating a non-idempotent write creates an extra record while preserving every checked field value. The verifier accepts all three task variants from two of five eligible generators: 6/15 constructed effects. A cardinality patch applied to checker copies after the census makes all six cells fail while preserving valid controls. In WorkArena, we prospectively rerun 23 extra-field candidates selected for earlier checker-PASS outcomes. Independent Table API readback confirms nondefault persisted values in 21, while all 23 receive PASS. Two requested strings are aliases of stored defaults. The 21 confirmed wrong effects span three form templates. These selected cases confirm wrong effects under the audit's protocol; they do not estimate a population rate. No other construction produces an independently confirmed false accept. Other checker-PASS cases are effect-correct degeneracies. We report zero-PASS families separately because retained evidence differs. In fixed intent-swap grids, the checkers return no PASS on 2,689 off-diagonal executions. This is a rejection census: 57 WorkArena cells use session-scoped evidence; the other 2,632 lack classified rejection causes and independent target ground truth. Each increment is specified before its own cells are scored. A supplement accompanies the OpenReview submission with the construction grammar, evidence, content-bound stage lineage and count reproducer.

Every Monday
Get next week’s papers.
Subscribe on Substack