Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files

Yupu Wang, Zhengyuan Jiang, Reachal Wang and Neil Zhenqiang Gong (Duke) introduce the package hallucination attack, in which a poisoned rule file such as AGENTS.md, CLAUDE.md or .cursorrules makes a coding agent import an attacker-controlled package instead of a legitimate dependency.
Ask this paper
Threat model. Rule files are shared on open platforms and marketplaces. An attacker injects a prompt into a benign rule file and publishes a malicious package to a registry such as PyPI, which is installed when the generated code runs.
PackHallu. An evolutionary optimizer rewrites the injected prompt using trajectory-level feedback from a surrogate agent and LLM-guided mutations.
Results. Across 3 coding benchmarks, 8 agent frameworks (including Claude Code, Cursor and OpenHands) and 13 backbone LLMs, PackHallu averages 79.29% syntactic and 67.68% deployable attack success, outperforming heuristic and optimization-based injection baselines.
Transfer. Prompts optimized only on an OpenHands surrogate still reach 71.55% and 62.99% on structurally different frameworks; 88% of victim settings exceed 50%.
Detection. State-of-the-art prompt injection detectors fail to reliably flag the malicious rule files.
Abstract
Modern agentic coding frameworks increasingly rely on community-shared rule files (e.g., this http URL or .cursorrules) to guide autonomous code generation, yet the security risks of this pipeline remain underexplored. To bridge this gap, we introduce the package hallucination attack, where an attacker injects malicious prompts into benign rule files to induce coding agents to replace legitimate dependencies with attacker-controlled packages. To obtain effective malicious prompts injected into rule files, we propose PackHallu, an evolutionary optimization framework that iteratively rewrites these injected prompts using trajectory-level feedback and LLM-guided mutations. Evaluations across multiple benchmarks, LLMs, and agent frameworks show that PackHallu achieves high attack success rates and strong transferability across diverse models and agent combinations. Our findings demonstrate that coding agents are vulnerable to package hallucination attacks, highlighting the urgent need for stronger security safeguards in autonomous coding systems.