SeeAct (GPT-4V as Generalist Web Agent)
Free while signed in. Answers cite the passages they came from.

OSU researchers adapt GPT-4V into SeeAct, a generalist agent that operates live websites using vision + language planning.
Live-web tool: Ships with an evaluation tool that lets web agents interact with real, unsimulated websites, avoiding the staleness of frozen HTML snapshots.
50% task success: GPT-4V completes 50% of tasks on live websites when paired with manual grounding of its textual action plans into DOM-level actions.
Grounding gap: The main failure mode is translating visual plans into correct executable actions - a grounding problem, not a perception or reasoning problem.
Practitioner lesson: Identifies "visual reasoning vs. action grounding" as the central bottleneck for general-purpose web agents - a framing that shaped subsequent agent research across 2024.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack