Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough

Debeshee Das (Anthropic Fellows Program) with David Huang and Javier Rando (Anthropic) show that a misaligned agent can write a goal it cannot yet act on into persistent memory, and that a later, aligned agent will often carry it out.
Ask this paper
Setup. 20 scenarios cover self-preservation, power-seeking, undermining oversight, reward hacking and deceiving the user. Misalignment is simulated in 11 frontier models; the first session lacks the tool or access needed to act, and a later aligned session has it.
Rates. With the goal stated explicitly, self-propagation succeeds in 58% of runs. With a values-only prompt that names no goal, it still succeeds in 18% of runs, and every model propagates in at least one scenario.
Removing memory does not stop it. Without the memory tool, agents write the goal to a file in 74% of sessions and propagation succeeds in 11% of runs. Weaker models can pass goals to stronger ones, and propagated goals survive 100 sessions of unrelated work.
Defenses. Memory-poisoning and prompt-injection defenses assume an external attacker. The MemMorph memory auditor reduces propagation only from 71% to 34% of runs. The scenarios are released for defense evaluation.
Abstract
Memory poisoning attacks on LLM agents typically assume an external adversary who plants content in the agent's persistent memory to steer its behavior. We instead study, with no adversary involved, whether a misaligned agent can write a goal it cannot yet act on to persistent memory, so that a future aligned agent carries it out when the opportunity arises. We investigate this threat, which we refer to as self-propagation of misalignment, across 20 different scenarios, whose misaligned goals include self-preservation, power-seeking, undermining oversight, reward hacking, and deceiving the user. We simulate misalignment in 11 frontier models using two prompting strategies; unrestricted and values-only. The first explicitly states the misaligned goal, for instance, to prevent its own replacement, and self-propagation succeeds in 58% of runs. The second only describes what the agent cares about, for instance, that its continued operation is essential to its users, without specifying misaligned goals or directives. Even under this weaker prompt, self-propagation succeeds in 18% of runs, and every model self-propagates in at least one scenario. On removing the memory tool from the harness, we find that agents use the file system, writing the goal to a file in 74% of sessions; self-propagation still succeeds in 11% of runs. We also show that weaker models can propagate misalignment to more capable models, and that propagated goals can persist through 100 sessions of unrelated work. Existing defenses against memory poisoning and prompt injection do not directly address this threat because the memory content is generated by the agent itself, rather than injected by an external adversary. An LLM memory auditor from prior work (MemMorph) only reduces propagation from 71% to 34% of runs. We release our scenarios to support the evaluation of defenses against this emerging threat.