🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 29, 2026
Agents · Evaluation

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

First page
The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge
The curator’s take

Alexander Gill, Kenneth Marino, Ana Marasović and colleagues at the University of Utah (EMNLP 2026 Findings) introduce KNOWS, a benchmark of browser tasks where the agent must research a topic and then produce a document, presentation or spreadsheet.

Ask this paper

Key points
01

Task design. Each open-ended, long-horizon task ends in a produced artifact and requires retrieval, synthesis, task decomposition and visual layout work inside real program interfaces. A written rubric and protocol govern task authoring.

02

Evaluators. Each task has a program combining deterministic checks with LLM judgments; evaluators agree with experts at Cohen's kappa 0.64 and 82% pairwise accuracy.

03

Results. The best frontier computer-use agent fully succeeds on under 3% of tasks, while partial-success scores for the best baseline range from 35% to 70%.

04

Visual steps break artifacts. Failures on visual and layout steps make artifacts unusable even when agents complete more than half of the other evaluation steps.

Abstract

Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such workflows demand reasoning and synthesis, decomposition of complex tasks, as well as visual and spatial understanding. To study agents on workflows like these, we introduce KNOWS, a benchmark of open-ended, complex, browser-based tasks that jointly evaluate these capabilities, with each task culminating in a produced artifact. To write tasks, we develop a task design rubric and a protocol for ensuring that tasks meet the requirements. Each task is paired with an evaluator, a program that combines deterministic checks with LLM judgments to balance the richness, reliability, and automation tradeoff inherent to agent evaluation. We evaluate and analyze frontier computer-use agents and browser-based harnesses. They achieve moderate scores on partial-success metrics, but the best performer fully succeeds in fewer than 3% of our complex, long-horizon tasks. Failures on visual steps render the resulting artifacts unusable, even when agents complete more than 50% of other evaluation steps. Our results expose limitations of current agents acting as end-to-end assistants, and call for progress on tool use, visual understanding, and long-horizon reasoning.

Every Monday
Get next week’s papers.
Subscribe on Substack