🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 18, 2026
Agents

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

First page
ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI
The curator’s take

Jaehyun Nam, Jinsung Yoon and colleagues at Google Cloud AI Research and the University of Waterloo present ScientistTwo, a multi-agent framework that takes a research problem, establishes baselines, forms hypotheses and runs an end-to-end discovery cycle without human intervention.

Ask this paper

Key points
01

Seed ideas target stated limitations. Idea generation is anchored to resolving limitations in existing work rather than sampling novel directions, which is what gives the loop a verifiable starting baseline.

02

Evaluation runs subset before full set. Ideas are screened on a data subset and only promoted to full experiments if they survive, which is the cost control that makes the cycle affordable.

03

Automated ablations feed refinement. The system runs its own ablation studies to attribute gains and rewrites the idea from the result, rather than reporting the first configuration that worked.

04

A simulated rebuttal engine closes the loop. Manuscript drafting simulates peer review and meta-review, and findings are revised from those reviews before the paper is finalised.

05

Benchmarked against accepted ICLR, ICML and NeurIPS papers. The reported solutions outperform the human state-of-the-art models on those problems and score higher average ratings than the human-authored papers under automated AI reviewers, which is a measurement of the AI reviewers as much as of the system.

Abstract

Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and venture into unexplored territory. The ultimate vision for AI in science is problem-driven autonomous research: given a fundamental challenge by a human expert, the AI independently navigates the scientific landscape, uncovers theoretical and empirical bottlenecks, and systematically expands the frontier of knowledge. In this paper, we introduce ScientistTwo, a fully autonomous multi-agent framework designed to realize this vision. Specifically, ScientistTwo takes an initial problem as input, establishes state-of-the-art baselines, formulates novel hypotheses, and coordinates specialized agents to orchestrate an end-to-end discovery cycle without human intervention. Moreover, the framework rigorously conducts experiments using diverse datasets and metrics, refines methodologies through automated ablation studies, and validates research findings via a closed-loop simulated peer-review rebuttal engine. To evaluate ScientistTwo's capabilities against the highest standards of human scientific achievement, we benchmark it across papers accepted at top-tier conferences such as ICLR, ICML, and NeurIPS. As a result, ScientistTwo autonomously generates expert-level, publishable papers and fully verified, executable codebases. Its solutions consistently outperform human state-of-the-art models, and achieve higher average review ratings than human-authored papers under automated AI review agents. These results show that ScientistTwo is not merely an assistive tool but an autonomous scientific pioneer capable of pushing the frontiers of human discovery. Project website: https://scientist-two.github.io/

Every Monday
Get next week’s papers.
Subscribe on Substack