🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 6, 2026
Agents

LEAP: Learning Efficient Action Proposals For LLM Agents

First page
LEAP: Learning Efficient Action Proposals For LLM Agents
The curator’s take

Zhen Xu, Qizheng Zhang, Gerry Wan, Shang Zhu and Ce Zhang (University of Chicago, Stanford, Together AI) present LEAP, which trains a 0.6B draft model to propose the next agent action so a larger target can verify it in parallel, cutting end-to-end agent latency.

Ask this paper

Key points
01

Latency framework. Speedup depends on hit rate, draft cost and how many steps a round can commit, so a larger, more accurate draft can be slower. With measured hit rates and draft costs the framework's predictions fall within 0.10 of measurement on 19 of 21 configurations.

02

Training. Every action the target commits is a free training pair. A LoRA adapter trained without reproducing reasoning text raises a 0.6B draft's agreement with Qwen3-32B from 2-25% to 64-75%; an online variant starts near sequential speed and approaches the offline draft.

03

Results. Across four benchmarks, two target families and three draft sizes on one shared GPU, the 0.6B draft gives 1.10x to 1.63x whole-task speedup. Task success rises in 8 of 16 comparisons, falls in 5 and is unchanged in 3, never by more than three tasks.

04

Comparison. On tau-bench, a reasoning-free copy of the target agrees on 42.5% of contexts against 63.7% for the trained 0.6B draft.

Abstract

LLM agents are known to be slow in rollouts. An agent completes a task one step at a time. At each step, it reasons and then chooses an action to execute. The next step and action cannot start until the previous one has finished. Speculative decoding accelerates the rollouts at the reason phase by drafting and verifying the inference tokens. Recent works have also started to apply similar ideas at the action phase. These works use off-the-shelf models, usually large, to draft action proposals for target model to verify. Large drafters match the target more often but take longer to propose, while small off-the-shelf models are fast but rarely make the same decision as the target. We ask a more general question: what determines the end-to-end speedup of action speculation? To answer it, we develop a latency framework for the speculative round. The framework compares what a round gains with what it costs. The gain depends on how well the drafter predicts the target and on how many steps the task can take before it ends. The cost comes from drafting, from waiting for target verification and from executing tools. Guided by the framework, we introduce LEAP (Learning Efficient Action Proposals) which keeps the drafter small and makes it accurate by training it on the target actions sequences. With a small 0.6B model, LEAP agrees with the target on most decisions and makes agents up to 60% faster in end-to-end wall clock time, with no systematic change in task success. Across various datasets, target models and draft models, the framework accounts for most of the measured speedups. We also show the draft model can be online trained with no prior trace collection and match the performance of offline training, making LEAP practical to deploy in the real world.

Every Monday
Get next week’s papers.
Subscribe on Substack