🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 21, 2026
Safety · Reinforcement Learning · Training

Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning

First page
Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
The curator’s take

Yuan, Kang, Liu, Choi, Iyer, Jiang and Jaques (University of Washington and Stanford) propose MoDA, an online RL post-training method that counters alignment-induced mode collapse by conditioning one shared policy on numbered roles that are rewarded for producing outputs distinct from each other.

Ask this paper

Key points
01

MARL framing. Each abstract numbered role acts as an agent in a cooperative-competitive game over one policy, so no hand-written personas or architecture changes are needed.

02

Quality gate. Diversity reward is granted only to responses that clear a prompt-adaptive quality threshold, which prevents the policy from gaining diversity by producing worse answers.

03

Diversity gains. On Infinite-Chat held-out prompts SBERT diversity rises 265% over Qwen3-8B; against the strongest DivPO baseline SBERT diversity goes from 0.274 to 0.482 and E-Vendi from 2.86 to 4.4.

04

No capability cost. Average general-capability pass@1 improves by 10.3% over the Qwen3-8B baseline and by 7.0% over DivPO across seven general tasks and four ideation and creative-writing tasks.

Abstract

A notable byproduct of LLM alignment training is mode collapse: the progressive loss of output diversity that narrows a model's expressivity at inference time. This degradation is especially limiting for applications requiring open-ended exploration and pluralistic perspectives, such as scientific ideation and creative writing. We present MoDA (Mode-conditioned Diversity Alignment), an online post-training RL algorithm that jointly optimizes generation quality and diversity, inspired by the coordination perspective in multi-agent reinforcement learning (MARL). MoDA trains a single shared LLM policy conditioned on abstract numbered roles, where each role acts as an agent competing to produce outputs distinct from the others. This formulation encourages mode-conditioned agents to explore complementary regions of the high-quality output space without requiring hand-crafted personas or architectural modifications. MoDA employs a prompt-adaptive quality gating mechanism that calibrates a reference quality threshold and grants diversity rewards only to responses that meet the threshold, preventing reward-hacking behaviors that compromise response quality. To study quality-diversity tradeoffs, we evaluate MoDA on a comprehensive suite of benchmarks spanning seven general capability tasks and four domain-specific diversity tasks in scientific ideation and creative writing. MoDA improves SBERT diversity by 265% on the Infinite-Chat held-out prompts, while increasing average general capability pass@1 by 10.3% over the Qwen3-8B baseline. Compared with the strongest DivPO baseline, MoDA improves SBERT diversity from 0.274 to 0.482 (+75.9%) and E-Vendi from 2.86 to 4.4 (+53.8%), while improving average general capability pass@1 by 7.0%. Overall, MoDA provides a drop-in alternative to standard post-training methods that preserves and expands the model's expressive output space while improving quality.

Every Monday
Get next week’s papers.
Subscribe on Substack