🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Agents

Threats in LLM-Powered AI Agents Workflows

Free while signed in. Answers cite the passages they came from.

First page
Threats in LLM-Powered AI Agents Workflows
The curator’s take

This work presents the first comprehensive, end-to-end threat model for LLM-powered agent ecosystems. As LLM agents gain the ability to orchestrate multi-step workflows and interact via protocols like MCP, ANP, and A2A, this paper surveys over 30 attack techniques spanning the entire stack, from input manipulation to inter-agent protocol exploits.

Key points
01

Four-part threat taxonomy: The authors categorize attacks into (1) Input Manipulation (e.g., prompt injections, multimodal adversarial inputs), (2) Model Compromise (e.g., composite backdoors, memory poisoning), (3) System & Privacy Attacks (e.g., retrieval poisoning, speculative side-channels), and (4) Protocol Vulnerabilities (e.g., MCP discovery spoofing, A2A prompt chaining).

02

Real-world attack success rates are high: Adaptive prompt injections bypass defenses in over 50% of cases, while attacks like Jailbreak Fuzzing and Composite Backdoors reach up to 100% ASR. These findings suggest that even well-aligned agents remain deeply vulnerable.

03

Protocol-layer threats are underexplored but critical: The paper exposes novel vulnerabilities in communication protocols (e.g., context hijacks in MCP, rogue agent registration in A2A), showing how subtle abuses in capability discovery or authentication can trigger cascading failures across agent networks.

04

Emerging risks in LLM-agent infrastructure: New modalities like vision-language agents (VLMs), evolutionary coding agents, and federated LLM training introduce their own threat vectors, including cross-modal jailbreaks, dynamic backdoor activation, and coordinated poisoning attacks.

05

Call for system-level defenses: The authors advocate for dynamic trust management, cryptographic provenance tracking, secure agentic web interfaces, and tamper-resistant memory as key defense directions. They also highlight the need for new benchmarks and anomaly detection methods tailored to agent workflows.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack