🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Aug 29 – Aug 29, 2026
Training · Data

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

First page
Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms
The curator’s take

Siye Wu and colleagues compare three ways to consolidate domain-expert RLVR models, Merge of task vectors, Mix RL of pooled datasets, and multi-teacher on-policy distillation, using shared experts and data across scales.

Ask this paper

Key points
01

Averages hide the real spread: Average performance differs by at most 1.4 points across the three paradigms, but the gap reaches 8.6 points on a single benchmark, with domain-level variation tracking task-vector geometry.

02

Each paradigm has a distinct binding constraint: Mix RL depends on domain mixture proportions, MOPD stays bounded by its teachers, and Merge compresses all expert updates into one set of weights.

03

No coverage gains anywhere: All three improve single-sample accuracy without measurable gains in solution coverage or losses in held-out capabilities. Fusion sharpens, it does not broaden.

04

A usable decision rule: Merge when experts exist and cheap fusion matters, Mix RL when training a unified model from scratch with tuned domain proportions, MOPD when preserving domain-specific gains outranks surpassing teachers.

05

Why it matters: Multi-capability post-training is where most labs are spending compute right now, and this is the first side-by-side under matched conditions.

Abstract

Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of large language models, but covering multiple capabilities often involves training separate domain experts and subsequently consolidating them. We organize three fusion paradigms by the artefacts they reuse: Merge combines expert task vectors, Mix RL pools their datasets, and multi-teacher on-policy distillation (MOPD) uses both. Because they have largely been studied in isolation, how they compare and how to choose among them remain unclear. We compare all three using shared experts and data across model scales and a multi-domain benchmark suite. Although their average performance differs by at most 1.4 points, the gap reaches 8.6 points on a single benchmark, with domain-level variation tracking cross-domain relations visible in task-vector geometry. Training dynamics expose distinct constraints: Mix RL depends on domain mixture proportions, MOPD remains bounded by its teachers, and Merge compresses all expert updates into one. All three improve single-sample accuracy without measurable gains in solution coverage or losses in held-out capabilities. These results yield a practical guideline: use Merge when experts already exist and cheap fusion is paramount; Mix RL when training a unified model without experts, with domain proportions adjusted for cross-domain transfer; and MOPD when preserving domain-specific gains matters more than surpassing teachers or minimizing end-to-end cost.

Every Monday
Get next week’s papers.
Subscribe on Substack