🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 4, 2026
Evaluation · Agents

EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks

First page
EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks
The curator’s take

Mukul Singh and colleagues at Microsoft introduce EmailBench, a self-contained benchmark of 206 enterprise email and productivity scenarios built on a typed email API and a synthetic Enron-style corpus, and show that agents complete most of their tool calls correctly while failing most tasks.

Ask this paper

Key points
01

Benchmark design. 206 scenarios in 16 categories over mail, calendar, contacts and tasks, graded by 258 executable state assertions plus 211 LLM rubrics. Topic selection was informed by aggregate task-intent telemetry from an internal prototype.

02

Headline result. The best configuration, Claude Sonnet 4.5, passes 33.5% of scenarios, followed by Claude Opus 4.1 at 29.6%. GPT-5 variants range from 20.9% to 28.2%.

03

Valid calls are not completed tasks. Sonnet's tool calls complete without an API error 99.7% of the time, yet scenario pass rates across all configurations stay between 20.9% and 33.5%.

04

Cross-domain tasks are the hard part. Sonnet passes 88.9% of folder tasks but 16.7% of calendar writes and 17.6% of orchestration tasks, and its pass rate falls from 37.7% on one-domain tasks to 14.3% on four-domain tasks.

05

Failure to act. GPT-5 and GPT-5.2 each return no tool call at all on 71 scenarios (34.5%), compared with 7 for Sonnet and 4 for Opus.

Abstract

Enterprise email agents must combine information retrieval, structured state changes, temporal reasoning, and multi-step coordination. Recent agent benchmarks include productivity tasks, but few center on typed email workflows in a self-contained environment. We introduce EmailBench, a benchmark of 206 email and productivity scenarios across 16 task categories. The benchmark couples a typed email API specification with provider-neutral naming, a deterministic synthetic Enron-inspired corpus, and a scenario suite whose topic selection was informed by aggregate task-intent telemetry from an interactive prototype. Its hybrid evaluation protocol combines 258 executable static assertions with 211 LLM rubrics. We evaluate eight LM configurations on a fixed single-user corpus. The best-performing configuration passes only 33.5% of scenarios despite 99.7% of its tool calls completing without an observed API failure, with pass rates varying substantially across task categories. This gap shows that valid tool execution is not equivalent to task completion. EmailBench provides a self-contained environment for end-to-end email-agent evaluation, with broader tool coverage, multi-persona testing, and repeated-run evaluation as future work areas.

Every Monday
Get next week’s papers.
Subscribe on Substack