🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 3, 2026
Agents · Training

RISED: RubrIcs for agentic multi-environment Selection and sElf-Distillation

First page
RISED: RubrIcs for agentic multi-environment Selection and sElf-Distillation
The curator’s take

Jingtan Wang and Bryan Kian Hsiang Low (NUS) with Sirajul Salekin, Young mok Jung, Javier Movellan and Manjot Bilkhu at Apple propose RISED, which uses LLM-judge rubric tags on rollouts to select training data and to supervise the policy when training one agent across several environments.

Ask this paper

Key points
01

Problem. In multi-environment RL, environments are learned at different rates, so all-success and all-failure groups appear in the same batch and give GRPO no group-relative signal, and scalar rewards say nothing about how rollouts in different environments relate.

02

Rubric profiles. An LLM judge tags each rollout with terms from a rubric vocabulary shared across environments; the resulting behavior profiles drive online selection of prompt groups that match the batch's overall behavior mix while avoiding overlap.

03

Self-distillation. Positive rubrics serve as privileged context for an on-policy self-distillation teacher that adds token-level supervision; negative rubrics steer later rollouts away from recurring failure modes.

04

Results. On ALFWorld, WebShop and DBBench, RISED has the highest mean pass rate on both backbones and ranks first or second in every environment; with Qwen2.5-3B the overall pass rate is 0.542 against 0.484 for GRPO-64 and 0.515 for GRPO-128.

Abstract

Training a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at the environment level or prioritize local reward-based signals, without explicitly considering relationships between current rollouts across environments for prompt-group selection. Meanwhile, as environments are learned at different rates, all-failure and all-success rollout groups can coexist within a batch, leaving those data without group-relative reward signals. Both challenges highlight limitations of relying solely on scalar rewards in multi-environment RL: they provide limited information about cross-environment relationships and no within-group reward contrast when rewards are identical. This motivates richer textual feedback, such as rubrics describing rollout behaviours, to guide learning. Beyond rubrics' usage as reward, we repurpose rubrics to guide both online data selection and policy supervision. An LLM judge tags each rollout using a predefined rubric vocabulary shared across environments. The resulting profiles guide the selection of data that aligns with the overall behavioural composition of the mixed-environment batch while limiting overlap with already-selected data. Available positive rubrics (describing desired behaviours) provide privileged context for an on-policy self-distillation teacher, supplying additional token-level supervision, while negative rubrics (describing undesired behaviours) guide subsequent rollout generation away from recurring failure modes. Together, these components form RISED. Across model backbones, RISED achieves the highest mean pass rate across environments and ranks first or second in every individual environment. Rubric-based analysis of RISED can further characterize the behavioural changes accompanying these gains.

Every Monday
Get next week’s papers.
Subscribe on Substack