🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 28, 2026
Evaluation · Agents

WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks

First page
WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks
The curator’s take

Yining Hua (Harvard, Agent Evaluation Science) and Levi Lian (Raycaster, Stanford) introduce WorkWorlds, an evaluation infrastructure that fixes an organization's state before any task is written, so benchmark construction cannot pre-select the evidence an agent needs.

Ask this paper

Key points
01

The flaw it targets. Many knowledge-work benchmarks choose each task's context together with or after the task. That places the relevant documents in the environment and completes part of the information-finding work an employee would normally do.

02

Design. A world fixes a revision, a date and an employee seat, and materializes only what that employee could access. Tasks are introduced afterward.

03

Implementation. A synthetic pharmaceutical company with 8 measured tasks across 6 employee seats, plus additional organizational worlds.

04

Measured effect of curation. Across 192 matched evaluations, task-level curation raised sufficient-evidence access from 72.8% to 90.4% (+17.6 points) and criterion pass from 68.0% to 76.7% (+8.7 points).

05

Where the gap comes from. Pass rate conditional on reaching the evidence was nearly unchanged, so most of the difference occurs before the agent finds the evidence.

Abstract

Many knowledge-work benchmarks are constructed around individual tasks, with the context needed for each task selected together with or after the task has been specified. This design measures performance on workplace-like tasks in an environment assembled for the task. When task specification guides which context is selected, the evaluation can encode task information into the environment and pre-complete part of the information-localization work that workplace performance normally requires. We introduce WorkWorlds, an evaluation infrastructure that separates organizational state from task specification. A world first fixes a revision, date, and employee seat and materializes the organizational state that employee can access; tasks are introduced only afterward. We implement WorkWorlds in a primary synthetic pharmaceutical company with 8 measured tasks across 6 employee seats, and construct additional organizational worlds. Across 192 matched evaluations, task-level curation increased evidence access by 17.6 percentage points, from 72.8% to 90.4%, and criterion pass by 8.7 points, from 68.0% to 76.7%, while pass conditional on evidence access remained nearly unchanged; most of the measured difference occurred before the agent reached sufficient evidence.

Every Monday
Get next week’s papers.
Subscribe on Substack