Structuring MoE Expert Selection for Agentic Reinforcement Learning

Bolian Li (Apple and Purdue) with Ting-Yao Hu, Cheng-Yu Hsieh, Oncel Tuzel, Raviteja Vemulapalli and colleagues at Apple study how mixture-of-experts routing relates to agentic behavior and control it during RL post-training.
Ask this paper
Observation. In off-the-shelf MoE models, expert routing overlaps more between turns where the agent performs similar operations, such as READ or UPDATE, than between turns doing different operations.
Problem. Standard RL leaves routing unconstrained, which the authors find limits both task success and inference efficiency.
Method. Turn-level expert selection is pushed to align with the agent's operation type, token-level selection is regularized for local consistency, and an entropy-gated control keeps training stable.
Result. The routing control framework improves success rate by more than 10 points on every evaluated benchmark.
Abstract
Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we observe expert selection exhibits a specialized structure that naturally aligns with agentic trajectories. Specifically, expert routing overlaps more between turns where the agent performs semantically similar operations (e.g., READ, UPDATE) than between turns with differing operations. However, standard RL algorithms ignore this specialization, allowing the MoE routing to go uncontrolled during training, which empirically limit task performance and inference efficiency. To address this, we introduce a hierarchical routing control framework for agentic tasks. We explicitly encourage turn-level expert selections to align with agentic operations while regularizing token-level expert selections to maintain local consistency. To resolve stability issues that arise during post-training with the proposed methods, we further introduce an entropy-gated control mechanism. Overall, our routing control framework achieves over 10-point improvements in success rate on all evaluated benchmarks. These results demonstrate that agentic trajectory structure provides an effective signal for optimizing MoE capacity during RL post-training.