🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 2, 2026
Agents

Schema: Discovering Unknown Environments via Agentic Program Induction

First page
Schema: Discovering Unknown Environments via Agentic Program Induction
The curator’s take

Guanning Zeng, Angjoo Kanazawa, Andrea Zanette, Haiwen Feng and colleagues at UC Berkeley, Carnegie Mellon and Impossible Research introduce Schema, an agent harness in which the LLM records what it learns about an unknown environment as executable programs instead of prose notes.

Ask this paper

Key points
01

Interactive program induction. The agent writes programs for the environment's state representation and transition rules, checks them against its interaction history, and plans inside them before acting.

02

Harness components. A persistent program workspace plus a small set of interfaces for testing programs against history, planning within them, and executing plans with step-by-step verification.

03

ARC-AGI-3. With the same base model, RHAE rises from 58.7% in the model's coding harness to 99.2% with Schema, and action efficiency improves across four model configurations.

04

Other environments. Schema solves all 21 public DiG-bench games with fewer failed attempts, and on MazeBench it reuses knowledge over tens of thousands of interactions to reach the median of the top-50 human players.

05

Why programs. Prose memory degrades under compaction; executable programs stay verifiable and can be reused for planning, which the ablations credit for the gains.

Abstract

Learning to complete tasks in unfamiliar environments with unknown rules remains a key challenge for LLM agents. Current LLM agents often record their discoveries in prose, which may not provide a compact, explicit account of how the environment works. Inspired by how scientists organize observations into testable, predictive theories, we introduce Schema, an agent harness that organizes learning and action through interactive program induction. The LLM agent decides what to investigate and how to act, expressing its evolving understanding of the environment as executable programs. The harness consists of a persistent program workspace and a small set of interfaces for checking these programs against the interaction history, planning within them, and executing plans under step-by-step verification. Schema raises ARC-AGI-3 RHAE from 58.7% to 99.2% with the same base model, solves 100% of the public DiG-bench games, and reaches the median performance of the top-50 human players on MazeBench. Extensive analysis shows the effectiveness of Schema in unknown mechanism discovery, and ablations confirm the contribution of each component.

Every Monday
Get next week’s papers.
Subscribe on Substack