ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning

Kun Feng, Yuchen Fang and colleagues at ShanghaiTech University and Ant Group introduce ARISE, an agentic RL framework that turns rollout evidence into paired rubrics and skills, retires criteria once mastered, and samples tasks by estimated capability, raising Qwen3.5-27B from 23.4% to 45.6% on SkillsBench.
Ask this paper
Problem. Fixed rubrics stop being informative once the agent satisfies them, and group-relative RL gives zero advantage when every rollout in a group gets the same reward, so mastered and impossible tasks waste rollout budget.
Rubric-skill pairs. Each observed capability gap becomes a rubric that rewards partial progress and a skill that guides action. Low rubric pass rates trigger skill use or refinement; consistently high rates retire the rubric.
Adaptive sampler. Task selection uses rubric-based estimates of capability within each task type, so training focuses on behaviors that are still weak for that kind of task.
Results. ARISE scores 45.6% on SkillsBench v1.1 and 50.6% on Terminal-Bench v2.1, up from 23.4% and 41.6% for its Qwen3.5-27B base, and above VADE, OnlineRubrics and RuscaRL (38.7% to 40.2% on SkillsBench).
Against larger models. The 27B model beats Qwen3.5-397B-A17B on SkillsBench (45.6% vs 30.3%) and stays close on Terminal-Bench (50.6% vs 51.3%).
Abstract
As a long-horizon agent improves through experience, previously observed weaknesses may recede while new limitations emerge, continually changing what it still needs to learn. Yet the learning process often remains tied to a static view of these needs: fixed behavioral criteria and training priorities can become misaligned with evolving agent capabilities, while sparse task-level feedback makes such misalignment more difficult to detect. Even when capability gaps are identified, rollouts from the current policy may repeatedly reproduce the same failures rather than explore better alternatives. To address this, we introduce Adaptive Rubric-Skill Co-Evolution (ARISE), a reinforcement learning framework that uses rollout evidence to continually adapt evaluation criteria, exploration guidance, and training priorities. Rubrics evolve to reward partial behavioral progress, while their paired skills are refined and selectively activated to guide exploration toward unresolved weaknesses. Alongside this co-evolution, capability-based adaptive sampling prioritizes tasks that target behaviors needing further improvement. Experiments on two challenging long-horizon agent benchmarks, SkillsBench and Terminal-Bench, demonstrate that ARISE successfully enhances both overall task performance and training efficiency. The project page is at https://foundation-model-research.github.io/ARISE .