🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 1, 2026
Agents · Training

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

First page
WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents
The curator’s take

Bo Mao, Tao Gui, Xipeng Qiu and colleagues at East China Normal University, Fudan and Shanghai Innovation Institute introduce WEFT, which scales tool-use post-training by evolving the whole interaction system (environment, task, harness and evaluator) rather than only the environments.

Ask this paper

Key points
01

Whole-system construction. Interaction systems are scaled across environment breadth, task complexity and interaction diversity, since environments alone do not give reliable learning signals.

02

Execution-driven self-evolution. Execution traces and state evidence are used to attribute failures to a component and revise it, with fresh rollouts checking each change.

03

Stable training. Prefix-preserving sampling keeps verified progress, atomic-turn credit assignment localizes the learning signal, and MegaMCP keeps isolated, recoverable state across concurrent rollouts on shared tool services.

04

Results. WEFT-8B and WEFT-14B beat all matched-size environment-scaling baselines on BFCL V4, tau2-Bench and Claw-Eval; WEFT-14B improves over Agent-World-14B by 6.41, 2.23 and 12.27 points, and WEFT-35B-A3B extends gains to Toolathlon-Verified and AutomationBench.

Abstract

Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee commensurate gains in model performance, because reliable learning signals depend on coherent interactions among all components of the agentic interaction system. To address this problem, we introduce WEFT (Whole-system Evolution For Tool-use Post-training), which couples scalable agentic interaction system construction, execution-driven self-evolution, and stable post-training. WEFT scales agentic interaction system construction across environment breadth, task complexity, and interaction diversity. Execution-driven self-evolution iteratively uses execution traces and state evidence to attribute failures and revise the responsible components, with fresh rollouts evaluating the changes and providing evidence for subsequent evolution rounds. For stable post-training at scale, WEFT addresses both optimization and execution reliability: prefix-preserving sampling retains verified progress and atomic-turn credit assignment localizes learning signals, while MegaMCP maintains isolated, recoverable state across concurrent rollouts over shared tool services. Extensive experiments across various models and benchmarks demonstrate the effectiveness of WEFT for tool-use post-training. WEFT-8B and WEFT-14B outperform all evaluated matched-size environment-scaling baselines on BFCL V4, $τ^2$-Bench, and Claw-Eval. In particular, WEFT-14B improves over Agent-World-14B by 6.41, 2.23, and 12.27 percentage points. WEFT-35B-A3B further extends these gains to more challenging long-horizon workflow benchmarks, including Toolathlon-Verified and AutomationBench.

Every Monday
Get next week’s papers.
Subscribe on Substack