Loopjacking: Hijacking Human-in-the-Loop Approval

Arun Kumar (independent) names Loopjacking: a person approves what they see as operation A while the agent runtime uses that approval for a different operation B, and reproduces it in released agent frameworks.
Ask this paper
Two variants. In representation attacks B is already encoded but hidden or misrendered at approval time; in state-substitution attacks the approver sees the correct A and mutable workflow state later swaps in B.
Reproductions. Post-approval substitution reproduced in seven Agno AgentOS releases up to 3.0.9 and 12 versions of a LangGraph Agent Server composition up to 0.14.0; representation mismatch in OpenClaw 2026.2.23, rejected in 2026.2.24.
Negative control. OpenAI Agents SDK 0.22.0 and 0.22.2 preserve exact per-call binding and reject a mutated B.
Mitigation. Rendering the complete canonical operation at approval and comparing it exactly at use time, or blocking unauthorized pending-state mutation, stops the tested attacks. The paper does not estimate prevalence.
Abstract
Human approval is often treated as the last security boundary before an agent executes a consequential operation. That boundary is only meaningful if the operation presented for review is the operation later authorized or released. We call failures of this binding Loopjacking: a human approves what they understand as operation A, while the implementation uses that decision for a materially different operation B. We distinguish two variants. In a representation-based attack, B is already encoded but omitted or misrepresented at approval time; in a post-approval state-substitution attack, the human sees the correct A and mutable workflow state later replaces it with B. We evaluate a purposive set of released agent products. We reproduce post-approval substitution in seven tested Agno AgentOS releases ending at 3.0.9 and in 12 tested versions of a conditional in-memory LangGraph Agent Server composition ending at 0.14.0. We reproduce representation mismatch in OpenClaw 2026.2.23 and its rejection in 2026.2.24. OpenAI Agents SDK 0.22.0 and 0.22.2 provide a negative control: serialized continuation preserves exact per-call binding and rejects mutated B. These results do not estimate ecosystem prevalence. They show that complete canonical approval rendering and exact use-time comparison, or preventing unauthorized pending-state mutation, block the tested attacks while preserving legitimate execution. We separate this contribution from established work on misleading dialogs, session smuggling, action binding, and authorization continuity.