Recursive Self-Improvement through Multi-Agent Self-Supervision

Hyunin Lee (UC Berkeley, intern at Sakana AI), Yujin Tang, Jinglue Xu, Matei Zaharia and colleagues propose Multi-Agent Self-Supervision (MASS), a recursive self-improvement method for non-verifiable tasks where the model is its own optimizer and evaluator.
Ask this paper
Loop. MASS alternates evolutionary search over multi-agent workflows with supervised fine-tuning on the model's own trajectories. A single base model proposes, executes and evaluates workflows under structural guardrails.
Search space. Workflows are computational graphs of distinct roles and information routing, and the search finds which roles and routes work best for a task.
Results. Over two MASS cycles with Qwen3.6-27B, performance per output token rises 1.2 to 1.6x on four open-ended public benchmarks.
Data efficiency. A student trained on multi-agent traces outperforms a single-agent student trained on 1.4x more tokens.
Recursion. Because the improved model is a better optimizer and evaluator, each cycle produces a stronger starting point for the next.
Abstract
Recursive self-improvement (RSI) of a model on non-verifiable tasks, such as open-ended research, faces a supervision bottleneck when its outputs exceed what even human experts can reliably assess, leaving the model itself (optimizee) as the best available optimizer and evaluator. However, a single model instance struggles to critique and improve its own complex reasoning under this homogeneous loop. To address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories. Guided by early findings that multi-agent topologies excel at complex reasoning, MASS prompts a single base model to iteratively propose, execute, and self-evaluate multi-agent workflows. Through an evolutionary search constrained by structural guardrails, the model optimizes these computational-graph-like orchestrations, discovering the most effective distinct roles and information routing for a given task. Over two MASS cycles with Qwen3.6-27B, the model achieves 1.2-1.6x higher performance per output tokens on four open-ended public benchmarks. Because the improved model subsequently acts as a better optimizer and evaluator, this alternating framework enables a continuous, recursive bootstrapping of the model's capabilities. Moreover, multi-agent traces are also more training-efficient: a student trained on them outperforms a single-agent student trained on 1.4x more training tokens. These findings suggest that jointly learning orchestration and bounded subagent execution from multi-agent trajectories can provide an effective signal for RSI.