Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

Zihao Zhu and Baoyuan Wu (CUHK-Shenzhen) with Siwei Lyu (Buffalo) and Adel Bibi (Oxford) (NeurIPS 2026) introduce skill cascading attacks, where a harmful objective is split across several agent skills that each look benign when inspected alone.
Ask this paper
Threat model. In a prescription-review example, one skill weakens signals of discontinued medications, a second downgrades interactions tied to them, and a third suppresses the resulting low-priority alert, so a severe warning never reaches the physician.
SkillCascade. An automated multi-agent red-teaming framework, released with SkillCascade-Bench of 213 validated cascading test cases across agent systems and domains.
Attack success. Across three agent systems (including OpenClaw, Claude Code and Codex) and eight backbones, cascades induce harmful behavior at an average rate of 89.4%. Per-backbone rates range from 76.5% to 97.1%; flagship models resist best and smaller variants in each family are more vulnerable.
Defenses miss it. Cascades evade per-skill scanners, joint-skill scanners and runtime defenses, with a mean detection-evasion rate of 88.5%, because each edit stays inside declared capabilities.
Host matters little. Differences across agent hosts stay under 6 points for any backbone.
Abstract
A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-party capabilities, but the openness of this skill ecosystem also opens up a new attack surface. Prior work has focused on vulnerabilities within individual skills, but little attention has been paid to risks that arise from interactions across skills. In this paper, we introduce skill cascading attacks, a threat paradigm in which a malicious objective is distributed across multiple skills so that each modification looks benign in isolation, yet their combined execution is harmful. For instance, in a prescription-review pipeline, the first skill weakens signals of recently discontinued medications in the extracted history, the second downgrades the severity of any drug interaction tied to them, and the third suppresses the resulting low-priority alert in the final summary, so that a severe drug-interaction warning silently disappears before reaching the physician. To systematically study this safety blind spot, we develop SkillCascade, an automated multi-agent red-teaming framework, and release SkillCascade-Bench, a benchmark of 213 validated cascading test cases across multiple agent systems and domains. Across representative agents (e.g., OpenClaw, Claude Code, Codex) and LLM backbones, cascaded interactions reliably induce harmful behaviors while evading existing per-skill scanners and runtime monitors. Our findings highlight a gap between component-level integrity and system-level safety, and call for defenses that reason over cross-skill interactions rather than individual skills in isolation.