🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 2 – Sep 2, 2026
Agents · Reinforcement Learning · Evaluation

ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

First page
ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning
The curator’s take

Fanrui Zhang and a large Alibaba-affiliated team propose ARISE-RL, a co-evolutionary loop in which a task and rubric Generator and a reasoning Solver train each other, replacing the verifiable gold answer that open-ended agentic RL does not have.

Ask this paper

Key points
01

Rubrics grounded in tool observations. The Generator writes criteria against real tool outputs rather than abstract quality dimensions, and is rewarded for producing intermediate-difficulty tasks that sit at the Solver's current capability boundary.

02

Selective self-distillation. Reward-Gated Self-Evolution Distillation distills a memory-augmented variant of the policy back into itself only when the memory actually improved empirical reward, which avoids the usual failure of imitating noisy guidance.

03

A new evaluation surface. ECR-Bench is an expert-calibrated rubric benchmark covering single-tool deep research and multi-tool travel planning, aimed at the open-ended tasks where binary verifiers do not apply.

04

Stability is the headline claim. The authors emphasize robust and stable state-of-the-art results across benchmarks, targeting the brittle-reward regime where rollout contrast is weak and group-based policy learning normally degrades.

Abstract

Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold answers and scalable rubrics. Moreover, even near the model's capability boundary, long-horizon open-ended agentic tasks often yield brittle and unstable rewards, resulting in weak or noisy rollout contrast that obscures fine-grained optimization signals for group-based policy learning. To address these challenges, we propose ARISE-RL, a novel full-cycle self-evolution framework that couples a task/rubric Generator and a reasoning Solver through rubric-mediated co-evolution. The Generator grounds tool-related rubric criteria in real tool observations and is rewarded for producing valid, intermediate-difficulty tasks aligned with the Solver's evolving capability boundary. The Solver, in turn, learns from fine-grained rubric satisfaction signals through multi-step reasoning and tool use. We further introduce Reward-Gated Self-Evolution Distillation (RG-SED), which selectively distills a memory-augmented variant of the same policy back into itself only when the memory yields empirical reward improvement, thereby reducing distribution mismatch and avoiding blind imitation of noisy guidance. Finally, to support rigorous evaluation, we present ECR-Bench, an expert-calibrated rubric benchmark suite covering single-tool deep research and multi-tool travel planning. Extensive experiments demonstrate that ARISE-RL consistently achieves robust and stable overall state-of-the-art performance across all evaluated benchmarks.

Every Monday
Get next week’s papers.
Subscribe on Substack