🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 10, 2026
Agents

Recursive Self-Improvement through Multi-Agent Self-Supervision

First page
Recursive Self-Improvement through Multi-Agent Self-Supervision
The curator’s take

Hyunin Lee (UC Berkeley, intern at Sakana AI), Yujin Tang, Jinglue Xu, Matei Zaharia and colleagues propose Multi-Agent Self-Supervision (MASS), a recursive self-improvement method for non-verifiable tasks where the model is its own optimizer and evaluator.

Ask this paper

Key points
01

Loop. MASS alternates evolutionary search over multi-agent workflows with supervised fine-tuning on the model's own trajectories. A single base model proposes, executes and evaluates workflows under structural guardrails.

02

Search space. Workflows are computational graphs of distinct roles and information routing, and the search finds which roles and routes work best for a task.

03

Results. Over two MASS cycles with Qwen3.6-27B, performance per output token rises 1.2 to 1.6x on four open-ended public benchmarks.

04

Data efficiency. A student trained on multi-agent traces outperforms a single-agent student trained on 1.4x more tokens.

05

Recursion. Because the improved model is a better optimizer and evaluator, each cycle produces a stronger starting point for the next.

Abstract

Recursive self-improvement (RSI) of a model on non-verifiable tasks, such as open-ended research, faces a supervision bottleneck when its outputs exceed what even human experts can reliably assess, leaving the model itself (optimizee) as the best available optimizer and evaluator. However, a single model instance struggles to critique and improve its own complex reasoning under this homogeneous loop. To address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories. Guided by early findings that multi-agent topologies excel at complex reasoning, MASS prompts a single base model to iteratively propose, execute, and self-evaluate multi-agent workflows. Through an evolutionary search constrained by structural guardrails, the model optimizes these computational-graph-like orchestrations, discovering the most effective distinct roles and information routing for a given task. Over two MASS cycles with Qwen3.6-27B, the model achieves 1.2-1.6x higher performance per output tokens on four open-ended public benchmarks. Because the improved model subsequently acts as a better optimizer and evaluator, this alternating framework enables a continuous, recursive bootstrapping of the model's capabilities. Moreover, multi-agent traces are also more training-efficient: a student trained on them outperforms a single-agent student trained on 1.4x more training tokens. These findings suggest that jointly learning orchestration and bounded subagent execution from multi-agent trajectories can provide an effective signal for RSI.

Every Monday
Get next week’s papers.
Subscribe on Substack