🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 20, 2026
Agents

Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents

First page
Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents
The curator’s take

Grace Chang Yuan, Pranav Rajpurkar and colleagues at MIT and Harvard Medical School study agents that manage a full emergency-department shift and introduce Asclepius, a harness that rewrites its own operating manual between shifts.

Ask this paper

Key points
01

Execution gap: On the Clinical Environment Simulator, agents usually reach the correct diagnosis but fail to deliver complete and timely critical actions.

02

Three failure modes: The gap comes from instruction-adherence drift, treatment incompleteness and a severity-equity gap in timeliness, each measured as a per-trace counter.

03

Harness design: Asclepius combines a self-evolving harness updated from trace feedback, an external clinical skills library for high-stakes regimens, and three isolated subagents that divide per-turn decisions across the patient queue.

04

Results: On held-out batches it improves critical-action correctness by 22% (p = 0.024) while keeping diagnostic accuracy, with gains consistent across five judges. Large reductions appear only when all three components act together.

Abstract

LLM agents are predominantly benchmarked on short, single-task trajectories, yet real deployments run for hours under contention, surfacing a different class of failures. We use the Clinical Environment Simulator (CES), in which an agent manages an entire emergency-department shift under continuous time and resource pressure, as a testbed: long-horizon execution failures manifest measurably in a single rollout under structured, multi-dimensional grading. On CES, current agents reach the correct diagnosis in most cases yet fail to deliver complete and timely critical actions, revealing an execution gap. We attribute this gap to three long-horizon failure modes, each operationalized as a per-trace counter: instruction-adherence drift, treatment incompleteness, and a severity-equity gap in timeliness. We then introduce Asclepius, an adaptive agent scaffolding with a self-evolving harness that rewrites the operating manual between shifts from trace-level feedback, an externalized clinical skills library for high-stakes regimen knowledge, and three isolated subagents that partition per-turn decisions across the patient queue. On held-out batches never observed during harness evolution, Asclepius improves critical-action correctness by 22% (p = 0.024) over a strong baseline agent framework while preserving diagnostic accuracy, with consistent gains across five LLM judges from three model families; on the full ten-batch set, improvements reach 25% on critical actions and 13% on timeliness. The three failure modes form a coupled bottleneck: decisive reductions appear only when all three components act together.

Every Monday
Get next week’s papers.
Subscribe on Substack