My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning

Yihua Zhu, Qianying Liu, Weixu Qiao and colleagues at Alibaba with Kyoto University propose FAULT, which converts an agent's own natural-language diagnosis of its errors into step-level credit that is anchored to the terminal reward in agentic RL.
Ask this paper
Problem. With terminal rewards, rollout groups where every sample has the same outcome give GRPO no signal, and a failed trajectory penalizes every step equally rather than the steps that caused the failure.
Method. A self-diagnoser names the erroneous steps, FAULT checks the diagnosis against trajectory evidence, and per-error-type costs are learned online from recent outcomes. The policy and the diagnoser are trained together.
Signal coverage. On ALFWorld, FAULT extracts a usable learning signal from 95% of rollout groups, against 41% for GRPO and 72% for GiGPO.
Results. Across 1.7B and 4B Qwen3 backbones, gains over raw-initialized GRPO and GiGPO are 17.1 to 41.7 points on ALFWorld and 8.7 to 21.0 in WebShop score. The margin over SEED, the strongest baseline, is smaller (6.7 and 3.0 points on ALFWorld), and on Search-based QA FAULT ties or trails SEED slightly.
Where it helps. The benefit is concentrated in long-horizon tasks where errors and their consequences are many steps apart.
Abstract
Agentic reinforcement learning (RL) has emerged as a powerful approach for training large language model agents on multi-step tasks, yet reliance on terminal outcome rewards creates two credit-assignment problems, particularly in long-horizon tasks. First, same-outcome rollout groups provide no learning signal from terminal rewards. Second, terminal rewards provide only trajectory-wide feedback, making it difficult to identify which decisions caused a failure. Recent work supplements terminal rewards with finer-grained information from trajectory analysis, such as natural-language reflections on intermediate decisions and errors. However, natural-language diagnoses are difficult to use directly for credit assignment: their error claims may be unreliable, and they do not quantify how much each error should affect learning. We propose Self-Diagnosis-guided Terminal Credit Redistribution (FAULT), which turns diagnosed errors into explicit step-level credit anchored by terminal outcomes. FAULT checks diagnostic evidence and learns relative error costs from task outcomes. During training, the policy and self-diagnoser co-evolve, while error costs are updated online from recent outcomes. On ALFWorld, FAULT recovers learning signals from same-outcome groups, reaching 95% signal coverage versus 41% for GRPO and 72% for GiGPO, while better localizing credit to specific error steps. Across two model scales, FAULT delivers strong. improvements on the long-horizon ALFWorld and WebShop tasks while remaining competitive on short-horizon Search-based QA.