🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Evaluation

ScienceAgentBench

First page
ScienceAgentBench
Paper summary

a new benchmark to rigorously assess agents built for scientific workflows; after testing it on open-weight and proprietary LLMs, the best-performing agent can only solve 32.4% of the tasks independently and 34.3% with expert-provided knowledge.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack