VISTA: A Visual Harness for Reasoning in an Interactive World

Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu and Kaiming He at MIT introduce VISTA, a visual harness that lets a general-purpose multimodal model act in interactive environments from image observations, with a lossless visual memory it can search and rearrange while it reasons.
Ask this paper
Design. The model sees raw screenshots instead of text grids, every past frame is kept unchanged in memory, and the model can retrieve old frames and reorganize its visual input on demand. There is no program synthesis and no hand-built world model.
ARC-AGI-3 result. With Claude Opus 5.0, VISTA scores a perfect 100.00 Relative Human Action Efficiency on all 25 public games, against 40.68 for the same model under the official implementation. It uses 7,302 actions in total, 57.4% fewer than the 17,135 used by first-time human players.
Without program synthesis. Concurrent systems that reach 99 to 100 on ARC-AGI-3 represent the screen as numeric grids and reason through executable code models; VISTA is the first reported system at that level that does neither.
Ablation path. Moving the official implementation from text grids to images, then adding longer limits, continuous conversation and persistent GUIDE.md and WORKING.md notes, raises RHAE at each step above the 40.68 official baseline.
Transfer. The same harness beats minimal-harness baselines on GameWorld, AI GameStore and BabyVision with little adaptation; on BabyVision accuracy rises from 41.0% to 63.2% because the model can inspect regions of a single image.
Abstract
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.