Prime Agent

Prime Intellect released an open-source harness built for long-horizon work, and what persists between runs sets it apart. Most harnesses reset everything except the files on disk, which caps how much a system can compound.
Ask this paper
The model programs its own context: A persistent IPython REPL lets the model process its context programmatically instead of reading a flat transcript, so filtering, aggregating, and re-deriving state become code the model writes rather than tokens it re-reads.
A Continual Harness carries the rest: Histories, memories, skills, prompts, and subagent specifications persist across trajectories. Improvements accumulate across runs instead of being rebuilt from scratch each time the agent starts.
The jump on ARC-AGI-3 is large: Holding the model class fixed, RHAE Best@1 moves from 30% to 95.5%. It also matches or beats native harnesses on long-context coding, GPU kernel generation, and autonomous nanoGPT speedruns.
Why it matters: This is a working reference implementation of the compounding-harness idea rather than a paper describing one, and it is open source. If you have been reading about self-improving harnesses and wanted something to run, start here. The state hierarchy diagram alone is a useful map of what belongs in the prompt and what belongs in managed storage.
Abstract
Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for long-horizon evaluation and coding-agent workflows. A persistent IPython REPL follows the Recursive Language Model abstraction for programmatic context processing and test-time compute, while Continual Harness preserves histories, memories, skills, prompts, and subagent specifications across trajectories. Recursive subagents coordinate through direct agent-to-agent communication, and the Agents View lets humans inspect and manage daemon-backed sessions. Prime Agent standardizes execution, recovery, verification, and resource accounting while leaving strategy construction to the model. This low-friction, expressive membrane prevents harness failures from becoming model failures and pushes measurement toward the model's true maximal underlying capability. Prime Agent raises ARC-AGI-3 RHAE Best@1 from 30% to 95.5% and matches or exceeds native and popular harnesses across long-context coding, GPU-kernel generation, emulator construction, and autonomous nanoGPT speedruns. On Factorio, we find refinement allows for continuous technology progression and dedicated subagents enable parallelized work. Code is available at https://github.com/PrimeIntellect-ai/prime-agent.