🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 6, 2026
Evaluation · Agents

OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine

First page
OpenGameEval: Benchmarking Agentic Programming and Exploration in a Stateful Game Engine
The curator’s take

Eray Turkel and colleagues at Roblox release OpenGameEval, a benchmark that runs language-model agents inside reproducible Roblox Studio sessions and separates observation tools from editing tools so exploration behavior can be measured directly.

Ask this paper

Key points
01

Benchmark. 84 human-curated core tasks, an eight-tool action space, and executable checks on both the edited scene and a simulated play session. 13 frontier models were run with 16 attempts per task.

02

Hard for current models. The best model solves 51.7% of tasks on one attempt and 39.4% on five of five attempts, and no model solves six of the tasks in any attempt.

03

Frontier models differ by task type. Five models finish close together overall but solve different tasks; splitting by kind of work spreads them by 5.0 points on script-authoring tasks and 12.5 points on scene-change tasks.

04

Exploration predicts success. Holding task and model fixed, a run that inspects every object the reference solution touches before acting passes 13.4 points more often on scene-only tasks and 9.8 points more often on script-only tasks than a run that inspects none.

05

Release. Tasks, place files, annotations, a Studio plugin and a leaderboard under the MIT license; a shorter version is at the NeurIPS 2026 Workshop on Evaluation of Interactive Agents.

Abstract

We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable checks, both on the edited scene and in a simulated play session. Most agentic coding benchmarks require exploration but score only final task success. OpenGameEval separates observation tools from editing tools in its eight-tool action space, so exploration can be measured directly. We measure the pass rates and exploration behavior of 13 frontier models on 84 human-curated core tasks, with 16 attempts per task. The tasks are hard for current models. The best model solves 51.7% of tasks on a single attempt and 39.4% five times out of five, and no tested model solves six of the tasks. Models at the frontier reach similar pass rates by solving different tasks: splitting tasks by the kind of work they require spreads the top five by 5.0pp on script-authoring tasks and 12.5pp on scene-change tasks. Exploration behavior predicts whether a run succeeds. Holding task and model fixed, a run that inspects every object a reference solution touches before acting on it passes 13.4pp more often than a run that inspects none of them on scene-only tasks, and 9.8pp more often on script-only tasks. We release the task suite, its place files, the per-task annotations, a plugin that runs the tasks inside Roblox Studio, and an updated leaderboard under the MIT license at https://github.com/Roblox/open-game-eval.

Every Monday
Get next week’s papers.
Subscribe on Substack