PRO-LONG
Free while signed in. Answers cite the passages they came from.

Long-horizon tasks force a harness to decide what to save from a long stream of observations and how to load it back into context, and richer summaries usually make the exact detail you need harder to retrieve. PRO-LONG sidesteps this tradeoff with programmatic memory.
Keep everything, search it: Rather than compressing history into bespoke memory, PRO-LONG keeps a complete, structured interaction log and leans on coding-agent tooling to search that history on demand, so no observation is discarded up front.
A minimal framework: The design is deliberately lightweight, avoiding hand-built memory harnesses and instead treating the full log as a searchable artifact the agent queries when it needs a specific past detail.
Strong, cheaper results: On the full ARC-AGI-3 public game set, it improves over a base coding agent by an average of 18.0 points across frontier models, and matches or exceeds specialized state-of-the-art harnesses at up to 76.1% pass@1 while using 4.2 to 5.8 times fewer tokens.
Why it matters: It shows that for exploratory, long-horizon settings, a simple searchable log can beat elaborate memory engineering on both accuracy and cost, which is a practical recipe teams can adopt now.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack