FrontierChallenge: Evaluating Scientific Workflow Completion

Liangcai Su and a sixteen-author team release FrontierChallenge, a cross-domain benchmark of end-to-end scientific workflows where the best of twelve frontier models across three agent scaffolds completes only a fifth of tasks.
Ask this paper
Deliverable bundles, not final answers: Each of the 97 released tasks (of 300 total) fixes inputs and specifies a bundle of required scientific deliverables, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science and electrochemistry.
20.6% pass rate at the top: The best configuration completed 20 of 97 tasks under the full-completion criterion. Average Score captures partial progress separately, which keeps the benchmark from bottoming out at zero and losing signal.
Three scaffolds, twelve models: Varying the scaffold as well as the model is the right axis here, because scientific workflow completion is exactly where harness design should dominate raw model quality.
Why it matters: Headroom this large is rare in a 2026 benchmark. Full-completion scoring on multi-deliverable workflows is a much harder target than the single-answer scientific QA benchmarks that are already saturating.
Abstract
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.