🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 2, 2026
Agents · Evaluation

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

First page
OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
The curator’s take

Dingyuan Dai, Heli Qi, Lei Liu and a large team led from Tsinghua (Jie Tang, Juanzi Li) with CMU, Yale, Waterloo and UC collaborators introduce OSWorld-Science, a benchmark of 146 tasks in which computer-use agents must operate real scientific software and produce checkable artifacts.

Ask this paper

Key points
01

Task coverage. Workflows include molecular drawing and retrosynthesis (e.g. ASKCOS), pathology image analysis, statistical computing and physical simulation, proposed by domain experts and refined through human-AI co-design.

02

Artifact-based scoring. Task-specific evaluators inspect application state and outputs such as molecular structures, segmentation masks, plots and numbers, and give partial credit.

03

Results across 12 VLMs. Claude Fable 5.1 scores highest at 73.7% mean task score ($6.5 per task), followed by Claude Opus 5 at 57.9% and GPT-6 Astra at 57.4%; GPT-5.6 sol reaches 50.1% at $2.79 and GPT-5.6 luna 22.6% at $0.23.

04

Analysis. The paper studies the effect of language, reasoning effort and context length, and documents failure modes such as declaring a task done after partial work without checking.

Abstract

Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorld-Science, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain. The benchmark contains 12 VLMs and 146 high-quality tasks across several scientific domains and software configurations, covering workflows such as molecular drawing and retrosynthesis, pathology image analysis, statistical computing, and physical simulation. Tasks are developed through expert proposals and iterative human--AI co-design, with selection guided by scientific value and difficulty. Task-specific execution-based evaluators inspect application states and generated artifacts, including molecular structures, segmentation masks, plots, and numerical results, and award partial credit for incomplete outcomes. Our special harness integrates model adapters, interaction-loop control, and trajectory logging to support comparisons of models and interaction strategies. Our results show that current state-of-the-art VLMs with a strong harness still face challenges in addressing key questions in the scientific domains. We also analyze the benchmarking results across multi-linguistics, reasoning efforts, context length and other factors and derive several important conclusions and directions to assist future development. Overall, we provide an integrated framework connecting expert-defined scientific goals to verifiable software outcomes, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.

Every Monday
Get next week’s papers.
Subscribe on Substack