🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 7, 2026
Agents · Safety

MLLMs Fail to Refuse when Using Tools Agentically

First page
MLLMs Fail to Refuse when Using Tools Agentically
The curator’s take

Rikiya Takehi (MIT, during an internship at NVIDIA) with Ryo Hachiuma, Shaona Ghosh and colleagues at NVIDIA show that giving multimodal LLMs tools makes them refuse harmful requests less often; the paper is accepted at NeurIPS 2026.

Ask this paper

Key points
01

Finding. On MM-SafetyBench, HoliSafe and VLSBench, every tested agentic MLLM refuses less often with tools than without, with a relative refusal-failure increase of up to 68.7% and 17.7% on average.

02

Models affected. The drop appears in agent-tuned open models such as AdaReasoner, general open models such as Qwen3.5-122B-A10B, Claude Opus 4.6 and 4.7, and Gemini Agentic Vision. Claude Opus 4.6, the best no-tool refuser, goes from 13.6 to 18.1 refusal-failure rate.

03

Two causes. Context dilution, where tool outputs bury the harmful intent of the original request, and safety focus displacement, where the model prioritizes describing tool observations over the safety decision.

04

Mitigation. Re-inserting the original request and image before the final response partly restores refusals, and the authors recommend evaluating safety in the tool-use setting, not only in plain chat.

Abstract

Agentic multimodal large language models (MLLMs) have recently pushed the frontier of visual reasoning by calling tools such as zooming and tagging. Despite the recent strong success of agentic MLLMs, this work uncovers a critical safety failure in the tool-use paradigm: agentic tool-using MLLMs become less capable of refusing harmful requests. Our experiments confirm that, across three popular safety benchmarks, all the top open- and closed-weight MLLMs we test exhibit significantly lower safety in tool-using settings than in non-tool settings, with a relative refusal failure rate increase of up to 68.7%. Based on analysis of 100,000+ responses, including extended experiments, we also propose two possible reasons for this safety degradation.

Every Monday
Get next week’s papers.
Subscribe on Substack