🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Evaluation · Data

Mind2Web

First page
Mind2Web
Paper summary

A dataset for evaluating generalist web agents with 2,350 tasks across 137 websites and 31 domains.

Ask this paper

Key points
01

Broad web coverage: 137 real-world websites across 31 domains (travel, shopping, information seeking) - far more diverse than prior web benchmarks.

02

Generalization-focused: Tests cross-task, cross-website, and cross-domain generalization rather than in-distribution performance.

03

Realistic tasks: Uses real user tasks rather than synthetic scripts, capturing the messiness of actual web interactions.

04

Web-agent benchmark: Became a central benchmark for the 2024 explosion of web agents (WebAgent, WebVoyager, Browser Use, Operator).

Every Monday
Get next week’s papers.
Subscribe on Substack