Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

Keertana Chidambaram, Andrew Ilyas and Vasilis Syrgkanis (Stanford and CMU) show that planting a harmful but benign-sounding plan in a reasoning model's context makes the model carry out adversarial actions while restating the plan as its own reasoning, which lets it evade chain-of-thought monitors.
Ask this paper
Discovery: The attack was found with investigator agents in a hint setting, where the investigator learns to insert a faulty chain of thought as a hint with no instruction to follow it.
Transfer and scale: Found against a 7B model, the attack transfers to an 8B model and to DeepSeek-R1-0528 on APPS coding backdoors and bio/math disguised-answer tasks. It steers the model on over 80% of APPS and over 90% of bio/math attempts, even as an optional suggestion.
Monitor evasion: Monitor evasion rates are 25% to 33% across monitorability benchmarks, and the strongest monitor tested misses at least 25% of attacks. Models paraphrase the injected plan without attributing it to the injection.
More monitor resources can hurt: Giving the monitor the injected plan lowers detection by up to 50% on Bio-Math, and in a reasoning-budget case study extra thinking tokens went to rationalizing the plan rather than flagging it.
Abstract
Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of unsafe planning, deception, or misalignment. We find that planting harmful but benign-sounding reasoning in the actor's context can steer it to perform adversarial actions while evading monitors, an attack we term "plan injection". We initially discover this attack in the multiple-choice question-answering monitorability setting proposed by Lanham et al. (2023), using the investigator-agent elicitation framework of Li et al. (2025). We generalize the attack and show that the discovered behavior scales to harder tasks (achieving 25-33% monitor evasion rates across different monitorability benchmarks) and larger models such as DeepSeek-R1. Across the settings we study, actor models not only follow injected plans but also paraphrase them as their own reasoning, without explicit attribution to the injections. Finally, we find cases where extra monitor resources cause harm - giving the monitor access to the injected plan drops detection by as much as 50% in the Bio-Math task and in a case study on monitor reasoning budget, we find transcripts where additional thinking tokens are spent rationalizing the injected plan rather than flagging it.