IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation

Jiarui Liu (CMU, internship at Meta), Wen-tau Yih, Xin Luna Dong and colleagues at Meta Reality Labs introduce IdeaScientist, a three-role agent system trained with RL to generate research proposals by transferring mechanisms from analogous problems in other fields.
Ask this paper
Roles. Ideation is decomposed into gap finding, innovation and report writing. Each role is trained separately with RL: the gap finder identifies limitations in related work, the innovator draws solution intuitions from analogous settings, and the writer turns them into full proposals.
Svalbard Idea Vault. A corpus of 2.77M decomposed research ideas used for retrieval, training and temporally controlled evaluation.
Evaluation. Agents see only literature published before a cutoff date, and proposals are scored by how closely they match directions later explored in 15K human-authored papers.
Results. On Qwen3.6-27B, IdeaScientist beats the strongest open-source autoresearch baseline by 14.0%, mainly through gains in novelty, and outperforms Claude Code SDK with Claude-4.8-Opus and Codex SDK with GPT-5.4 by up to 5.9%.
Abstract
Despite rapid progress in automating scientific research, generating promising and well grounded research solutions remains a central challenge. We isolate research ideation as a standalone task and build our solution on the intuition that a challenge in one field can often be addressed by a mechanism that solved an analogous challenge in another. Accordingly, we introduce IdeaScientist, which decomposes ideation into gap finding, innovation, and report writing, and trains each role with reinforcement learning. These roles identify limitations in related work, draw solution intuitions from analogous problem settings, and develop those intuitions into complete research proposals. To facilitate discovery of insights across domains, we construct the Svalbard Idea Vault, a corpus of 2.77M decomposed research ideas for retrieval, training, and temporally controlled evaluation. Our evaluation restricts access to literature available before a cutoff date and assesses how closely proposed directions align with those later explored in 15K papers authored by human researchers. On Qwen3.6-27B, IdeaScientist outperforms the strongest open-source autoresearch baseline by 14.0%, driven mainly by gains in novelty. On this 27B open backbone, IdeaScientist even outperforms Claude Code SDK with Claude-4.8-Opus and Codex SDK with GPT-5.4, by up to 5.9%.