🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 23, 2026
Agents

Loopjacking: Hijacking Human-in-the-Loop Approval

First page
Loopjacking: Hijacking Human-in-the-Loop Approval
The curator’s take

Arun Kumar (independent) names Loopjacking: a person approves what they see as operation A while the agent runtime uses that approval for a different operation B, and reproduces it in released agent frameworks.

Ask this paper

Key points
01

Two variants. In representation attacks B is already encoded but hidden or misrendered at approval time; in state-substitution attacks the approver sees the correct A and mutable workflow state later swaps in B.

02

Reproductions. Post-approval substitution reproduced in seven Agno AgentOS releases up to 3.0.9 and 12 versions of a LangGraph Agent Server composition up to 0.14.0; representation mismatch in OpenClaw 2026.2.23, rejected in 2026.2.24.

03

Negative control. OpenAI Agents SDK 0.22.0 and 0.22.2 preserve exact per-call binding and reject a mutated B.

04

Mitigation. Rendering the complete canonical operation at approval and comparing it exactly at use time, or blocking unauthorized pending-state mutation, stops the tested attacks. The paper does not estimate prevalence.

Abstract

Human approval is often treated as the last security boundary before an agent executes a consequential operation. That boundary is only meaningful if the operation presented for review is the operation later authorized or released. We call failures of this binding Loopjacking: a human approves what they understand as operation A, while the implementation uses that decision for a materially different operation B. We distinguish two variants. In a representation-based attack, B is already encoded but omitted or misrepresented at approval time; in a post-approval state-substitution attack, the human sees the correct A and mutable workflow state later replaces it with B. We evaluate a purposive set of released agent products. We reproduce post-approval substitution in seven tested Agno AgentOS releases ending at 3.0.9 and in 12 tested versions of a conditional in-memory LangGraph Agent Server composition ending at 0.14.0. We reproduce representation mismatch in OpenClaw 2026.2.23 and its rejection in 2026.2.24. OpenAI Agents SDK 0.22.0 and 0.22.2 provide a negative control: serialized continuation preserves exact per-call binding and rejects mutated B. These results do not estimate ecosystem prevalence. They show that complete canonical approval rendering and exact use-time comparison, or preventing unauthorized pending-state mutation, block the tested attacks while preserving legitimate execution. We separate this contribution from established work on misleading dialogs, session smuggling, action binding, and authorization continuity.

Every Monday
Get next week’s papers.
Subscribe on Substack