🚀NEW LABGetting Started with Claude AgentsStart lab
Agents

SeeAct (GPT-4V as Generalist Web Agent)

First page
SeeAct (GPT-4V as Generalist Web Agent)
Paper summary

OSU researchers adapt GPT-4V into SeeAct, a generalist agent that operates live websites using vision + language planning.

Ask this paper

Key points
01

Live-web tool: Ships with an evaluation tool that lets web agents interact with real, unsimulated websites, avoiding the staleness of frozen HTML snapshots.

02

50% task success: GPT-4V completes 50% of tasks on live websites when paired with manual grounding of its textual action plans into DOM-level actions.

03

Grounding gap: The main failure mode is translating visual plans into correct executable actions - a grounding problem, not a perception or reasoning problem.

04

Practitioner lesson: Identifies "visual reasoning vs. action grounding" as the central bottleneck for general-purpose web agents - a framing that shaped subsequent agent research across 2024.

Every Monday
Get next week’s papers.
Subscribe on Substack