How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression

Xijie Gong, Tingxu Han, Lijie Hu and colleagues at MBZUAI (with Griffith and other universities) trace how agentic LLMs decide whether to call a tool or answer directly. The paper is accepted at NeurIPS 2026.
Ask this paper
Contrastive pairs. Long agentic prompts are reduced to 500 minimal pairs across Python, Java and C++ where one request verb sets the decision: an execution verb such as write triggers a call, an analysis verb such as discuss does not.
A tool-call vector. Activation patching localizes the decision to the layer-24 residual stream of Qwen3-8B, where patching recovers the tool call on 100% of corrupted prompts. The extracted vector is both necessary and sufficient.
Steering. Removing the vector suppresses 93.1% of baseline tool calls, and adding it induces calls on 70.0% of turns that originally answered in text; a norm-matched random vector changes 1.5% or less.
Mechanism. The scaffold sets a prior toward calling a tool, and analysis verbs suppress it through features signalling that tool use is unnecessary. MLPs supply 77.7% of the write in the key window.
Generality. The vector generalizes to multi-turn tau2-Bench trajectories and verb-free requests, and the same mechanism appears in seven models from the Qwen, Mistral and Granite families.
Abstract
Tool calling, invoking external tools on demand, is central to agentic LLMs, yet the mechanism that decides whether a model calls a tool or responds directly remains poorly understood. Agentic prompts are long and heavily scaffolded, combining role instructions, tool schemas, format templates, and the user's request across hundreds of tokens, creating a noisy, highly entangled context in which no single controllable variable for mechanistic analysis is obvious. To obtain such a variable, we propose a method that converts complex agentic prompts into minimal contrastive pairs in which a single request verb determines the tool-call decision: replacing an execution-verb (e.g., \textit{write}) with an analysis-verb (e.g., \textit{discuss}) reliably flips the decision, suggesting it is mediated by a compact internal state. We construct 500 such paired prompts across Python, Java, and C++ (300 for mechanistic analysis, 200 held out for evaluation). We trace the decision to a vector, $μ_Δ$, that is both causally necessary and sufficient and generalizes beyond the discovery prompts to native multi-turn $τ^2$-Bench trajectories and verb-free requests. Behavioral ablations show that the scaffold establishes a tool-call prior; Transcoder decomposition then reveals that analysis verbs suppress this prior through features signaling that tool use is unnecessary, whereas execution verbs largely leave it intact. Downstream scaffold-reading attention heads and MLP features read out the resulting state, and the same mechanism recurs across seven models from the Qwen, Mistral, and Granite families. Our code is available at https://github.com/XijieGo/MI4ToolCalling.