A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents

Yichun Feng (UCAS), Jiawei Wang (USTC) and Haozhe Sun (Meituan) propose SkillPivot, which updates an agent's natural-language skills from the point where a failed trajectory goes wrong instead of from the whole failure.
Ask this paper
Premise. Failed trajectories usually contain a useful prefix of evidence gathering followed by an erroneous suffix, and tool-use tasks often have several valid solution paths, so matching a single successful trajectory is the wrong target.
Deviation point. SkillPivot locates the step where productive work stops using execution validity, goal progress and action diversity.
Teacher continuation. A stronger teacher resumes from the same prefix under the same interaction history, with an inline reflector giving a hint after each tool step.
Localized updates. Contrasting the student's failed suffix with the teacher's successful suffix yields small, conditional, regression-checked skill edits that keep existing guidance that already works.
Evaluation. On ToolQA (six disjoint task groups), LogicBench and WildClawBench it outperforms competing skill-evolution methods across several agent models and produces compact updates that transfer.
Abstract
Large language model agents increasingly rely on natural-language skills to solve complex tool-use tasks. However, such tasks often admit multiple valid solution paths, making it inappropriate to improve skills by forcing failed trajectories to match a fixed successful trajectory. Moreover, failed trajectories are rarely entirely wrong: an agent may first collect useful evidence and make meaningful progress, but later deviate into an erroneous suffix. We therefore argue that skill self-evolution should identify where productive problem solving begins to break down, rather than reflect coarsely over the entire failure. Based on this insight, we propose SkillPivot, a deviation-point-guided framework for skill self-evolution. SkillPivot detects the transition from a useful prefix to an erroneous suffix using execution validity, goal progress, and action diversity. A stronger teacher then continues from the same prefix and produces a successful alternative under the same interaction history. By contrasting the student's failed suffix with the teacher's successful suffix, SkillPivot generates localized skill updates while preserving already effective guidance. Experiments on ToolQA, LogicBench, and WildClawBench show that SkillPivot consistently outperforms competing skill-evolution methods, improves multiple agent models, and produces compact, transferable skill updates.