🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 24, 2026
Agents

Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents

First page
Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents
The curator’s take

Kaijie Chen, Chenyu Fang, Peng Ye and colleagues at Tongji University, Shanghai AI Laboratory and Fudan propose Trace, which compiles noisy sparse-reward trajectories into short, state-conditioned procedures that an agent can execute and verify.

Ask this paper

Key points
01

Problem with raw experience. Full trajectories carry failures, loops and detours, while summaries leave out the state conditions and action dependencies needed to repeat a success.

02

Credit plus dependencies. Trace finds progress anchors from rewards and persistent state changes, propagates credit to valuable transitions, and estimates action prerequisites from success and failure across episodes.

03

Backward slicing. It traces each required fact back to the action that produced it, keeps the dependency-consistent chain and drops irrelevant loops.

04

Walkthrough format. Each walkthrough has entry conditions, ordered state-action-effect steps and completion and failure predicates, so it can resume from intermediate states and be checked programmatically.

05

Results. On J-TTL, WebShop and ScienceWorld with three open models, Trace beats eight test-time learning and memory baselines, improving average AUC by 30.0% and Final-3 by 40.5% over the strongest one with fewer inference tokens.

Abstract

Test-time self-evolving agents improve by reusing past experience, yet sparse-reward trajectories contain failures, loops, and detours, while summaries often omit the state conditions and action dependencies needed for execution. We study executable Walkthrough induction from sparse-reward trajectories: extracting compact, state-conditioned, and verifiable procedures. Our key observation is that delayed credit identifies actions associated with progress but cannot determine whether they produce facts required by later actions. We propose Trace, a credit-guided, dependency-grounded framework that compiles noisy trajectories into executable Walkthrough Memory. It detects progress anchors from rewards and persistent state changes, propagates credit to identify valuable transitions, and estimates action prerequisites from cross-episode success and failure evidence. Backward dependency slicing then traces required facts to their producers, extracting dependency-consistent action chains while removing irrelevant loops and detours. The resulting Walkthroughs encode entry conditions, ordered state--action--effect steps, and completion and failure predicates, supporting reuse, intermediate-state resumption, and programmatic verification. Experiments on J-TTL, WebShop, and ScienceWorld with three open-source LLMs show that Trace consistently outperforms eight test-time learning and memory baselines. Compared with the strongest baseline, it improves average AUC and Final-$3$ by $30.0%$ and $40.5%$, respectively, while using fewer inference tokens. These results show that long-horizon interaction benefits more from state-conditioned executable procedures than from complete trajectories or abstract summaries.

Every Monday
Get next week’s papers.
Subscribe on Substack