DAYJOB: A Benchmark for Long-Horizon Professional Work

Stephanie Finley, Liudas Panavas, Sushant Mehta, Edwin Chen and colleagues at Surge AI release DAYJOB, 130 healthcare and finance tasks written by working professionals, each estimated at 13 to 17 hours of human work and graded all-or-nothing against an expert rubric.
Ask this paper
Task design. 50 healthcare and 80 finance tasks start from a brief request, so the agent has to decide what is needed, which documents matter and whether the request's premise is true. Each runs as a containerized Harbor environment.
Grading. Expert rubrics have a median of 47.5 (healthcare) and 57.5 (finance) binary criteria; an agentic judge applies them to the delivered files, and an attempt passes only if every criterion is met.
Results. Across 30 configurations from 13 developers, the best, Claude Opus 5.5, passes 24.7% of healthcare and 23.9% of finance attempts. The median configuration passes 0.6% and 2.5%.
Failure pattern. Case studies show agents accepting premises that the record contradicts and carrying a wrong input through an otherwise consistent analysis, so the deliverable looks complete and is wrong.
Release. All healthcare tasks, 50 of 80 finance tasks, the harness and a leaderboard are public.
Abstract
Professional work often starts with a brief request that leaves the professional to work out what is needed, which documents matter, and whether the request's premise holds. We introduce DAYJOB, a benchmark of 130 tasks built by professionals in healthcare (50) and finance (80). The tasks are estimated to take a professional 13.6 hours on average in healthcare and 16.6 in finance. Each task is a containerized Harbor environment with an expert rubric of binary criteria (median 47.5 and 57.5 per task) that an agentic judge applies to the delivered files, and an attempt passes only if it meets every criterion. Across 30 model configurations from 13 developers, the strongest, Claude Opus 5.5, passes 24.7% of healthcare and 23.9% of finance attempts, and the median configuration passes 0.6% and 2.5%. In case studies, agents accept premises that the record contradicts and carry wrong inputs through otherwise consistent analyses. We release all healthcare tasks, 50 of the 80 finance tasks, the evaluation harness, and the leaderboard.