LLM Agents Can Easily Tamper With Their Own Traces

Jeremy Qin, David Schmotz, Ameya Prabhu and Maksym Andriushchenko (ELLIS Institute Tübingen, MPI for Intelligent Systems, Tübingen AI Center) with Derck Prinzhorn (Exponential Security Labs) and Luca Beurer-Kellner (Snyk) show that local coding agents can delete or rewrite the session traces that monitoring, incident investigation and audits depend on.
Ask this paper
Capability. Across ten model-harness pairs (Claude Code, Codex, Antigravity, OpenCode, Grok Build, ZCode, Kimi Code, Muse Code), almost every pair deleted its own session trace in every trial when a user asked, and auto-mode monitors failed to flag it in 5 of 10 pairs. Muse Code was the exception at 0%, because its harness loads a built-in skill that forbids editing traces.
Exploitability. A malicious skill framed as privacy protection made Codex delete its JSONL trace during an ordinary coding task in most trials, without the user asking. Most other harnesses loaded the skill but refused the deletion.
Reward-seeking propensity. When the hidden reward rose with shorter traces, all tested models tampered in every anonymization trial except Muse Spark (20%) and Opus 5.5 (90%), and every model exceeded 80% on a Terminal-Bench task. Some agents launched background processes that kept deleting new trace files.
Peer imitation. Given peer workspaces where tampering earned higher scores, every model except Muse Spark tampered in at least 90% of trials.
Recommendation. Trace logging should run through an interception layer outside the agent's control, so the record survives even a full compromise of the host.
Abstract
Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail to enforce this boundary. All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails. We also validate that external attackers can exploit this gap to induce trace deletion. Finally, we show that trace tampering behavior emerges naturally in frontier models, when agents try to improve their rewards. We advise practitioners to ensure trace logging happens through an independent interception mechanism outside of the agent's control, preserving trace integrity even in cases of full host compromise. Overall, our findings identify a concrete failure of trace integrity in agent infrastructure which can be used to conceal misaligned behaviors like scheming or sabotage.