🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 10, 2026
Reinforcement Learning · Architecture · Training

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

First page
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
The curator’s take

Xiaomi's LLM-Core team reports MiMo-V2.6, an omni-modal MoE family (Pro at 1.02T total / 42B active, Flash at 310B / 15B active) trained by scaling RL compute across batch size, environments and grader compute.

Ask this paper

Key points
01

Throughput. Asynchronous RL consumes 1,568 samples and 2.7-3.7B tokens per step at context lengths up to 1M on a hybrid sliding-window-attention architecture.

02

Environments. RL spans code, general, visual and cyber domains under a mixture of agent harnesses, supported by a unified trajectory representation, high-concurrency multi-framework rollout and decoupled control and data planes.

03

Grader compute. Groupwise agentic grading gives more accurate rewards on long-horizon tasks and pushes the model toward shorter, more token-efficient solutions.

04

Stability. The MoE router is frozen during RL to keep expert loads stable, and a multi-layer defense targets reward hacking.

05

Distillation. A MiMo-V2.6-Distill-Qwen-9B checkpoint raises SWE-bench Pro from 32.0 to 44.6 and AutomationBench from 5.0% to 30.3% over Qwen3.5-9B. Training dynamics, RL environments and the RL framework are open-sourced.

Abstract

Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute. Prior to RL, we conduct mid-training on a broad multimodal corpus to provide ample exploration space, and build a solid infrastructure on the pretrained hybrid-SWA architecture to support subsequent scale-up. We scale RL compute along three dimensions: (1) larger batches and higher throughput, with an asynchronous training that consumes 1,568 samples and 2.7-3.7B tokens per step at context lengths of up to 1M; (2) more diverse and complex environments, spanning code, general, visual, and cyber domains under a mixture of agent harnesses; and (3) more grader compute, via groupwise agentic grading that yields more accurate reward signals for long-horizon tasks and steers the model towards shorter, more token-efficient solutions. To keep training stable at scale, we freeze the MoE router and establish a multi-layer defense against reward hacking. We further build infrastructure for mixed-task agentic RL, including a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training-inference consistency. We open-source the training dynamics, RL environments, and RL framework to facilitate reproduction and further research on scaled RL and model self-improvement.

Every Monday
Get next week’s papers.
Subscribe on Substack