🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 16, 2026
Agents

PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress

First page
PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress
The curator’s take

Kevin Qinghong Lin, Pan Lu, Philip Torr, James Zou and colleagues (Oxford, Stanford, NUS) build PaperDoctor, an agent that gives authors pre-submission feedback in which every finding points to specific evidence and comes with a revision.

Ask this paper

Key points
01

Three layers: L1 screens writing, layout and references; L2 routes each claim to a typed verifier for code, theory, prior work or experiments; L3 reruns experiments in priority order under a compute budget.

02

Finding format: Each finding is a triple of observation, pointer to a sentence, equation or code line, and suggested revision.

03

Evaluation: On 30 in-progress papers it reaches 70.6% agreement with all-positive holistic scores; on 40 manuscripts across ML, natural and social sciences it gives more auditable feedback than human and agentic reviewers.

04

Reproduction: Selective reruns surface missing instructions, execution failures and numeric disagreements that a reading-only review cannot detect.

Abstract

Autoresearch agents are reshaping the research ecosystem, but they can also let flawed claims enter the literature at scale. Human advisors catch such issues in drafts through careful, traceable feedback, yet advisor-style assessment requires extensive manual effort and does not scale. To shift automated paper assessment from a judge to a diagnostician, we introduce PaperDoctor, an agent framework for pre-submission feedback with three key innovations. First, a holistic hierarchical framework evaluates writing, layout, references, code, theory, prior work, and experiments through three layers: L1 surface screening, L2 typed verifiers that route each claim to the appropriate evidence, and L3 reproducers that rerun experiments by priority. Second, each finding contains an observation, a pointer to specific evidence such as a sentence, equation, or code line, and a revision suggestion, making critiques auditable and actionable. Third, PaperDoctor selectively rebuilds and reruns experiments based on claim importance and compute budget, surfacing reproducibility gaps and quantitative limitations that are invisible from the manuscript alone. We evaluate PaperDoctor on 30 in-progress papers, yielding 70.6% agreement and all positive holistic scores, and on 40 manuscripts across machine learning, natural science, and social science, covering human- and AI-authored papers with code. Overall, PaperDoctor produces more auditable feedback than human and other agentic reviewers, pairs critiques with concrete suggestions by design, and complements dimensions often overlooked by human reviewers. We also develop an interactive interface that lets authors browse findings grounded in their paper. PaperDoctor reframes automated paper assessment as diagnosis rather than verdict, taking a concrete step toward AI advisors for more rigorous AI-assisted scientific discovery.

Every Monday
Get next week’s papers.
Subscribe on Substack