🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Agents

SeeAct (GPT-4V as Generalist Web Agent)

Free while signed in. Answers cite the passages they came from.

First page
SeeAct (GPT-4V as Generalist Web Agent)
The curator’s take

OSU researchers adapt GPT-4V into SeeAct, a generalist agent that operates live websites using vision + language planning.

Key points
01

Live-web tool: Ships with an evaluation tool that lets web agents interact with real, unsimulated websites, avoiding the staleness of frozen HTML snapshots.

02

50% task success: GPT-4V completes 50% of tasks on live websites when paired with manual grounding of its textual action plans into DOM-level actions.

03

Grounding gap: The main failure mode is translating visual plans into correct executable actions - a grounding problem, not a perception or reasoning problem.

04

Practitioner lesson: Identifies "visual reasoning vs. action grounding" as the central bottleneck for general-purpose web agents - a framing that shaped subsequent agent research across 2024.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack