🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 22, 2026
Agents · Evaluation

ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

First page
ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
The curator’s take

Kratika Bhagtani, Kusha Sridhar and colleagues introduce ERPBench, which evaluates screenshot-only computer-use agents on a live, reproducible ERP system and scores each task against values in its database.

Ask this paper

Key points
01

State-grounded scoring. Tasks are graded by the stored business records, which catches errors that never appear on screen.

02

Save is not correctness. Across six agents, strong general GUI skill does not transfer: some agents save the form in up to 85% of runs but write the correct value in as few as 3%.

03

Deployment harness. The authors also release a production harness that requires human approval before actions, which ERPBench runs autonomously.

Abstract

Computer-use agents that operate through screenshots and simulated actions are advancing rapidly, yet their evaluation remains anchored to general desktop and web tasks. Enterprise Resource Planning (ERP) systems run the finance, procurement, inventory, and customer operations of organizations worldwide, and pose distinct challenges for computer-use agents: dense interfaces, coordinated multi-step interactions, and errors that alter persistent business records rather than surfacing on screen. Existing enterprise benchmarks rely on proprietary platforms or on simulated approximations of such software. We introduce ERPBench, a benchmark that evaluates screenshot-only agents on a live and reproducible ERP system and scores each task against ground-truth values in its database. Beyond the benchmark, we present a production-grade harness that gates agent actions behind human approval for safe deployment, which ERPBench runs autonomously. Evaluating six closed and open-source agents, we demonstrate that strong general GUI performance does not transfer to enterprise reliability. Even when an agent reaches the right form and saves it, the stored record is often wrong: some agents save in up to 85% of runs but write the correct value in as few as 3%. We further characterize failure modes specific to enterprise workflows.

Every Monday
Get next week’s papers.
Subscribe on Substack