🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 23, 2026
Agents · Training

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

First page
DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
The curator’s take

Siyuan Liu, Yixin Cao and colleagues at Fudan University and the Meituan LongCat Team introduce DENSE, which turns an agent's own execution traces into structured feedback for a second attempt without needing outcome labels, verifiers or expert annotation.

Ask this paper

Key points
01

Nested shortcut trees. DENSE organizes a trace into nested subtasks, compresses redundant attempts, summarizes finished branches and expands unresolved ones, so the feedback links progress already made to requirements still open.

02

Issue reconciliation. Recovery evidence inside the trace is used to reconcile issues across levels of the tree, so problems the agent already fixed are not reported as open.

03

REFIT protocol. Feedback methods are compared from shared initial trajectories with outcomes hidden, and environments and model contexts are reset before the fresh attempt at the same task.

04

Results on Terminal-Bench 2.1. DENSE has the highest strict pass rate among non-privileged feedback methods across four recipient models, improving on the initial run by 7.12 to 15.64 points while using 19.0% to 43.6% fewer recipient tokens.

05

Ablations. GPT-5.5 ablations show that nested subtask analysis, shortcut construction and issue reconciliation each contribute to the gain.

Abstract

Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to distill these traces into reusable feedback without post-hoc outcome labels, drawing on their evidence of local progress, recovery, and unfinished requirements. We introduce DENSE (Distilling Evidence from Nested Subtask Executions), which organizes this evidence into evidence-grounded nested shortcut trees. DENSE compresses redundant attempts, reconciles issues across levels using recovery evidence, and summarizes completed branches while expanding unresolved ones, linking reusable progress to remaining obligations. We introduce REFIT, a source-paired protocol comparing feedback from shared initial trajectories under post-hoc outcome blindness, with environments and model contexts reset for fresh attempts at the same tasks. On Terminal-Bench 2.1, DENSE achieves the highest strict pass rate among tested non-privileged feedback methods across four recipient models. Relative to initial executions, strict pass rate improves by 7.12-15.64 pp, with 19.0-43.6% fewer observed recipient tokens in reruns. GPT-5.5 ablations support combining nested subtask analysis with shortcut construction and issue reconciliation. These findings point toward agent self-refinement through evidence-grounded trajectory reuse with less reliance on external supervision.

Every Monday
Get next week’s papers.
Subscribe on Substack