Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning

Chen Wu, Josh Passenger and Yin Song (AWS) trace how a stateless coding agent forms, carries and abandons knowledge across ARC-AGI-3 levels by following every belief it commits to a file. Accepted at the NeurIPS 2026 CL4FMAgents workshop.
Ask this paper
Setup. A frozen foundation model in a fixed harness plays ARC-AGI-3 by writing and running Python and shell scripts, keeping no state across turns except the files it writes, while the harness logs every action and observation.
Protocol. A measurement protocol traces each thought from the task where it forms to the task where it is corrected or abandoned, applied to seven evaluation runs with three backbones from two model families.
Scripts are not reused. Only 33 of 630 script references cross a task boundary because most scripts embed level-specific state. The agent rewrites general rules into new scripts and abandons 74% of the scripts written before a boundary.
Notes are append-only. Notes are never revised; the agent appends without removing earlier claims and resolves the resulting contradictions against the log, so it forgets selectively rather than catastrophically.
Costliest error. The most expensive failure is a hard-coded value carried into a level where it no longer holds.
Abstract
We study how a coding agent learns across a sequence of abstract reasoning tasks. The agent runs on a frozen foundation model inside a fixed harness and acts by writing and running Python and shell scripts. It retains no state across turns other than its written artifacts, so every thought it forms, carries, corrects or abandons leaves a trace, where a thought is any belief, rule or plan committed to a file. We let the agent play ARC-AGI-3, a set of interactive reasoning games that provide no instructions. Each game is a sequence of levels, and a strategy that clears one level can fail on the next, so every new level is in effect a new task. The agent records what it learns as Python scripts and text notes, while the harness keeps a complete log of every action and observation. Our contribution is a measurement protocol that traces each thought through these files, from the task where it forms to the task where it is corrected or abandoned, applied to seven evaluation runs with three backbones from two model families. Scripts written for one task are almost never called again in a later task (33 of 630 references cross a task boundary), because most scripts embed the state of the current level. Instead, the agent rewrites its knowledge into new scripts, keeping the general rules and dropping the level-specific details, and abandons 74% of the scripts it wrote before a boundary. The notes, which only the model reads, are never revised: the agent appends without removing earlier claims, and the contradictions that accumulate are settled against the log. Because the log preserves everything, the agent forgets selectively, not catastrophically. The most costly error is a hard-coded value carried into a task where it no longer holds. These findings come from the files the agent wrote, without access to the model, and constitute a white-box analysis of how a coding agent continually learns.