🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Aug 30 – Aug 30, 2026
Evaluation · Agents

FrontierChallenge: Evaluating Scientific Workflow Completion

First page
FrontierChallenge: Evaluating Scientific Workflow Completion
The curator’s take

Liangcai Su and a sixteen-author team release FrontierChallenge, a cross-domain benchmark of end-to-end scientific workflows where the best of twelve frontier models across three agent scaffolds completes only a fifth of tasks.

Ask this paper

Key points
01

Deliverable bundles, not final answers: Each of the 97 released tasks (of 300 total) fixes inputs and specifies a bundle of required scientific deliverables, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science and electrochemistry.

02

20.6% pass rate at the top: The best configuration completed 20 of 97 tasks under the full-completion criterion. Average Score captures partial progress separately, which keeps the benchmark from bottoming out at zero and losing signal.

03

Three scaffolds, twelve models: Varying the scaffold as well as the model is the right axis here, because scientific workflow completion is exactly where harness design should dominate raw model quality.

04

Why it matters: Headroom this large is rare in a 2026 benchmark. Full-completion scoring on multi-deliverable workflows is a much harder target than the single-answer scientific QA benchmarks that are already saturating.

Abstract

Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.

Every Monday
Get next week’s papers.
Subscribe on Substack