Codetta: High-Capacity, Keyless, and Undetectable Multi-Agent Collusion

Qi Pang, Virginia Smith and Wenting Zheng (Carnegie Mellon University) build Codetta, a steganographic protocol that lets two independently deployed LLM agents set up a shared key and exchange hidden messages over a monitored channel while their transcripts stay computationally indistinguishable from normal outputs.
Ask this paper
Threat model. Existing provably undetectable schemes either need the receiver to reproduce the sender's exact output distribution, carry very little data (under a small fraction of a bit per token for the prior asymmetric protocol), or rely on a pre-shared secret key. Codetta removes all three limits.
High-capacity channel. Both agents run a small public draft model over the transcript as a shared estimate of the channel, the sender samples in a way that preserves its output distribution, and an adaptive error-correcting code handles decoding errors. Capacity reaches up to 94 times the prior asymmetric protocol across three agent workloads and three sender models.
Keyless setup. A steganographic key exchange carries MLWE ciphertexts (the ML-KEM-512 assumptions) inside honest model samples. A 128-bit key takes about 80k visible tokens with certified failure probability of at most 4.1 x 10^-3, versus an estimated 2 million tokens for the prior theoretical construction.
Implication. Reading transcripts is not enough to detect collusion between agents. The authors point to auditing agents' actions and perturbing the communication channel instead.
Abstract
Multi-agent systems built on large language models (LLMs) are increasingly deployed in high-stakes settings such as finance, healthcare, and software engineering, where agents coordinate through natural-language messages. The same channels, however, let colluding agents exfiltrate confidential information or coordinate unauthorized actions, and steganography can hide such communication inside outputs that look ordinary to an auditor reading the transcript. Existing provably undetectable LLM steganography protocols are not suited to realistic deployments. High-capacity schemes assume a symmetric setting where the receiver can reproduce the sender's output distribution, the state-of-the-art protocol for asymmetric agents has very low capacity, and most approaches rely on a pre-shared secret key. We make the threat of undetectable agent collusion concrete with Codetta, a high-capacity steganographic protocol for independently deployed agents in realistic asymmetric settings. Codetta combines a shared public model that estimates the communication channel, a sampling mechanism that preserves the sender's output distribution, and an adaptive error-correcting code. It further removes the pre-shared key through a steganographic key exchange that lets independently deployed agents establish a shared key while keeping the transcript computationally indistinguishable from ordinary model outputs. Across three agent workloads and three sender models, Codetta achieves up to $94\times$ the capacity of the state-of-the-art asymmetric protocol, and its key exchange establishes a shared key with about 80k visible tokens at an empirically certified failure probability of at most $4.1\times 10^{-3}$. These results show that effectively undetectable collusion is becoming feasible between independently deployed agents, so auditing must go beyond inspecting communication transcripts.