LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles

Bingo Zhang and colleagues at Vera Praxis with Tencent, HKUST and CUHK introduce LongPuzzleBench, 114 levels of six visual puzzle games played only through GUI input, where a legal move can make a level unsolvable without any signal until several moves later.
Ask this paper
Design. 16 objectives chain up to ten levels under one time budget. A human needs a median of 204 GUI actions per objective and 1,106 for one Nut and Bolt level, and the game never announces a dead end.
Results. With native GUI actions only, GPT-6-Astra completes 91.7% of objectives, but seven of ten general-purpose agents solve nothing harder than Medium, the four GUI-specialized models complete none, and no agent finishes Bolt Unscrew Hard. A human solves all 16.
Code access is mixed. Allowing code raises scores for six of ten agents, but some agents act through the GUI in only 8.5% to 17.5% of steps and solve by transcribing the board and searching offline.
How agents fail. In unsolved attempts, only 15% of late moves paint new cells and 30% return to an earlier exact state. Retries reach the same dead end 79% of the time, and full history does not prevent it.
Rule knowledge is not the bottleneck. A rule card raises correctly stated rules from 43% to 88% but level completion only from 50% to 60%.
Abstract
GUI agents need long-horizon visual reasoning: they must interpret a changing interface while keeping a multi-step plan viable as earlier actions constrain later ones. Existing benchmarks evaluate grounding, computer use, and game play, but rarely test whether agents stay coherent across long chains of coupled decisions. Long-horizon visual puzzles expose this capability directly: a legal move that looks like progress can make the puzzle unsolvable, and the loss shows only several moves later. We introduce LongPuzzleBench, 114 levels in six puzzle games played through native GUI actions, where one objective can take a human over a thousand actions on persistent boards and dead ends go unannounced. With Native GUI Actions alone, the strongest agents solve most objectives, but success falls sharply on harder, longer boards: seven of ten general-purpose agents solve nothing harder than Medium, and none completes Bolt Unscrew Hard, which a human solves along with every other objective. Code Execution CUA does not close this gap, and its scores mix visual solving with algorithmic search. Controlled diagnostics trace these failures to one limitation that neither rules, state hints, nor failure memory removes: agents judge each move by the visible progress it makes, not by the future options it leaves.