🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 3, 2026
Agents

A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors

First page
A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors
The curator’s take

Pengxun Li and colleagues identify the lifecycle-hook update path as a new attack surface: agent harnesses trust hook configuration blindly, so a benign versioned plugin can be trojanized into running attacker commands at host privilege on events the LLM never sees.

Ask this paper

Key points
01

The threat model is narrow and realistic: the attacker controls only plugin metadata and lifecycle-hook configuration, not code, not the model, not the user's prompt.

02

Hooks fire outside the model's view: commands bound to session start, tool calls and file edits run with host privileges at moments the LLM never observes, so no amount of model-level alignment sees the attack.

03

HookPry compromises every harness tested: an open-source automated framework realizing ten attack objectives, across 25 harness-and-backend combinations in 1,000 end-to-end runs, compromising all seven evaluated harnesses with per-harness success up to 92.5%.

04

Existing defenses do not engage: Microsoft Defender has 0% recall and the union of three static defenses misses 47.5% of malicious artifacts.

05

Why it matters: hooks are now a standard extensibility mechanism in coding harnesses. This is a supply-chain attack aimed squarely at the layer developers have been adding fastest.

Abstract

Modern AI agent harnesses expose lifecycle hooks that bind shell commands to runtime events such as session start, tool calls, and file edits. These commands run with host privileges yet ship as lifecycle-hook configuration and may fire at times the LLM never observes. We identify the lifecycle-hook update path, which harnesses trust blindly, as a new attack surface. Under a supply-chain threat model in which an attacker controls only plugin metadata and lifecycle-hook configuration, a benign versioned plugin can be trojanized by an update that silently binds attacker-chosen commands to benign events, yielding malicious host-side behavior such as privilege escalation. We propose HookPry, an open-source and fully automated attack framework that systematically exploits this vulnerability across heterogeneous AI agent harnesses. HookPry realizes ten attack objectives; across 25 combinations of harnesses and backends in 1,000 end-to-end runs, it compromises all seven evaluated harnesses, with per-harness success rates reaching 92.5%. Representative defenses remain insufficient: Microsoft Defender has 0% recall, and the union of three static defenses misses 47.5% of malicious artifacts.

Every Monday
Get next week’s papers.
Subscribe on Substack