🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Agents · Evaluation · Data

Mind2Web

Free while signed in. Answers cite the passages they came from.

First page
Mind2Web
The curator’s take

A dataset for evaluating generalist web agents with 2,350 tasks across 137 websites and 31 domains.

Key points
01

Broad web coverage: 137 real-world websites across 31 domains (travel, shopping, information seeking) - far more diverse than prior web benchmarks.

02

Generalization-focused: Tests cross-task, cross-website, and cross-domain generalization rather than in-distribution performance.

03

Realistic tasks: Uses real user tasks rather than synthetic scripts, capturing the messiness of actual web interactions.

04

Web-agent benchmark: Became a central benchmark for the 2024 explosion of web agents (WebAgent, WebVoyager, Browser Use, Operator).

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack