🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 3, 2026
Reasoning · Code · Memory

VISTA: A Visual Harness for Reasoning in an Interactive World

First page
VISTA: A Visual Harness for Reasoning in an Interactive World
The curator’s take

Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu and Kaiming He at MIT introduce VISTA, a visual harness that lets a general-purpose multimodal model act in interactive environments from image observations, with a lossless visual memory it can search and rearrange while it reasons.

Ask this paper

Key points
01

Design. The model sees raw screenshots instead of text grids, every past frame is kept unchanged in memory, and the model can retrieve old frames and reorganize its visual input on demand. There is no program synthesis and no hand-built world model.

02

ARC-AGI-3 result. With Claude Opus 5.0, VISTA scores a perfect 100.00 Relative Human Action Efficiency on all 25 public games, against 40.68 for the same model under the official implementation. It uses 7,302 actions in total, 57.4% fewer than the 17,135 used by first-time human players.

03

Without program synthesis. Concurrent systems that reach 99 to 100 on ARC-AGI-3 represent the screen as numeric grids and reason through executable code models; VISTA is the first reported system at that level that does neither.

04

Ablation path. Moving the official implementation from text grids to images, then adding longer limits, continuous conversation and persistent GUIDE.md and WORKING.md notes, raises RHAE at each step above the 40.68 official baseline.

05

Transfer. The same harness beats minimal-harness baselines on GameWorld, AI GameStore and BabyVision with little adaptation; on BabyVision accuracy rises from 41.0% to 63.2% because the model can inspect regions of a single image.

Abstract

We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.

Every Monday
Get next week’s papers.
Subscribe on Substack