🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 8, 2026
Agents

From Evidence to Action: How Tool-Using Agents Fail

First page
From Evidence to Action: How Tool-Using Agents Fail
The curator’s take

Hongzhan Lin, Shidong Cao, Ziyang Luo and Wenhao Chai (Princeton) with Mong-Li Lee and Wynne Hsu (NUS) and authors at HKBU and Amazon Web Services study where tool-using agents break the chain from established evidence to action, and release SafeActBench.

Ask this paper

Key points
01

Benchmark. SafeActBench has 656 cases across six operational domains and five protocols that progress from static action judgment and investigated non-action to single-action and multi-action workflows.

02

Evidence ledger. A provenance-bound Evidence Ledger and deterministic trajectory evaluator record what information was established, when actions occurred and whether downstream dependencies were satisfied.

03

Static vs interactive. Across ten model-harness configurations, static accuracy of at least 95% coexists with interactive success of at most 52% on the same cases.

04

Where failures start. Agents often stop with incomplete investigation or act before required evidence is in place. Once evidence is complete, single actions succeed 93.2% to 100% of the time for nine of ten configurations; multi-action workflows add unresolved prerequisites.

05

Harness effect. Harness choice shifts results: DeepSeek gains 4.4 points with its own harness over Inspect, while GLM loses 6.8 points with ZCode, and two Qwen harnesses disagree on 23.5% of cases.

Abstract

Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand. We study where this evidence-to-action chain breaks as agents move from deciding whether to act to executing single actions and dependent workflows. Across ten model-harness configurations, strong static action assessment can coexist with much weaker interactive execution. Failures often begin before execution: agents stop with incomplete investigation or act before required evidence is established. Once required evidence is obtained, single-action execution is usually reliable, while multi-action workflows additionally expose unresolved prerequisites and incomplete execution. For this analysis, we introduce SafeActBench, comprising 656 cases across six operational domains and five protocols that progress from static action judgment and investigated non-action to single- and multi-action workflows. A provenance-bound Evidence Ledger and deterministic trajectory evaluator track what information was established, when actions occurred, and whether downstream dependencies were satisfied. These results show that failures arise not only from missing information, but also from how agents use established evidence when deciding and executing actions.

Every Monday
Get next week’s papers.
Subscribe on Substack