🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Code · Evaluation

Replicating ML Papers with Agents

First page
Replicating ML Papers with Agents
Paper summary

This work tests whether a coding agent can replicate a scientific ML paper from its materials alone, using a skill that turns each paper claim into a target with recorded evidence and gating completion on workspace evidence rather than the agent's final message. Across twelve runs over four papers, all twelve workspaces pass the completion gate and all 158 recorded targets are matched with report coverage. Yet repeated runs still differ in how papers are split into targets and in numerical fidelity, so completion becomes reproducible even when the path is not.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack