RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

Mohamed Dhouib, Sonia Vanier and colleagues at École polytechnique (LIX), with Elie Bursztein of Google DeepMind, introduce RAISED, a self-distillation defense that trains a tool-using agent to behave on injected contexts exactly as it behaves on clean ones.
Ask this paper
Diagnosis. Training-based defenses such as SFT, DPO and SecAlign++ shift the model's output distribution even on benign inputs, and on benign tool-use tasks they make the model skip steps that a tool output legitimately asks for.
Method. The model writes its own tool-use scenarios and injected variants, including cases where finishing the task requires following guidance from a tool output. The student is then trained to match the teacher's clean-context next-token distribution on both the clean and the injected copy of each trajectory.
AgentDojo results. On Gemma, attack success drops from 39.5% to 1.0% while utility under attack rises from 64.0% to 80.5%, and benign utility falls by only 1 point. PromptGuard2 on the same setup reaches 16.1% attack success but cuts utility under attack to 42.9%.
Limitation. Because RAISED preserves the base model's clean behavior instead of improving it, it also preserves that model's existing errors.
Abstract
Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We further identify a failure mode of these defenses: On benign tool-use tasks, the model refrains from a step needed to finish an authorized task, particularly when that step is indicated by a tool output. To address these limitations, we introduce RAISED (Robust Attack Invariance through Self-Distillation), a training framework that combines self-generation and self-distillation. The model first generates its own tool-use scenarios, with an emphasis on cases where task completion requires acting on legitimate guidance from tool outputs. Then, through self-distillation, the student is trained to match the teacher's clean-context behavior on both clean and injected variants of the same trajectory. RAISED substantially reduces the attack success rate of prompt injections in tool responses while, unlike prior training-based defenses, preserving utility on both agentic and general-purpose benchmarks.