Silent Failures in Agent-Tool Interaction: An Audit of ToolUniverse

Shreya Gopalan, Devansh Singh and Sundaraparipurnan Narayanan (AI Tech Ethics) audit 15 scientific tools in the ToolUniverse environment for silent failures, where a tool call appears to succeed but returns incomplete data and neither agent nor user is told.
Ask this paper
Definition. A silent failure is a successful-looking invocation whose information or functionality is partly or fully missing, with no error signal.
Method. LLM-based candidate discovery and automated tests, followed by manual validation, organized around 7 failure loci along the agent-to-tool chain.
Findings. 91 validated failures, most often missing data or fields and inconsistent search, filtering or ranking.
Where they originate. 51 failures sit in the API layer and 25 in the wrapper layer, upstream of the agent, and they propagate into scientific outputs that look valid.
Proposal. A notion of contextual reliability, with mechanisms to test, disclose, monitor and measure such failures across the pipeline.
Abstract
Agentic AI systems are increasingly adopting automated pipelines that integrate multiple tools. While prior research and benchmarks have studied about task success and task completion of these agentic systems, the research about agent to tool interaction, specifically in biology agentic workflow is limited. This study investigates specific failures in agent to tool interaction where a tool invocation appears successful, some or all of the information or functionality from the tool via API/ wrapper is incomplete or missing and there are no communications / notifications to the user or the agent about such missing information. We call this a silent failures as the user or the agents are not aware that such failure has occurred. For the purposes of this study we developed an audit mechanism to identify such silent failures in Agent to tool interaction, by examining 15 scientific tools (and their associated API documentation and tool documentations) integrated within ToolUniverse environment (ToolUniverse serves as our experimental environment rather than the object of the study itself). We structure our study around 7 failure locus characterising where the failure occurs in the chain. We observed 91 failures (manually validated post LLM based candidate discovery and automated testing), most frequent of them being missing data or fields and inconsistencies in search, filtering or ranking criteria. Most of the 91 failures occurred in API layer (51) or wrapper layer (25), with a potential of silent failure amplification downstream. The results show that silent failures originate upstream of the event and propagate downstream into apparently valid scientific outputs. We propose a concept of contextual reliability to handle such failures and suggest mechanisms for testing, disclosing, monitoring, and measuring such failures across the agent-tool interaction pipeline.