🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Agents

MedBrowseComp

First page
MedBrowseComp
Paper summary

MedBrowseComp is a new benchmark designed to evaluate LLM agents’ ability to perform complex, multi-hop medical fact-finding by browsing real-world, domain-specific web resources. Testing over 1,000 clinically grounded questions, the benchmark reveals major capability gaps in current models, with top systems achieving only 50% accuracy and GUI-based agents performing even worse.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack