Stealing Reasoning Traces
Free while signed in. Answers cite the passages they came from.

Frontier providers hide chain-of-thought and hand the client an encrypted block instead, which the client returns with every subsequent request. This work identifies an architectural flaw in that design and turns it into a scalable extraction attack across three providers.
The blocks are interchangeable: Encrypted reasoning blocks are fully compatible across sessions, users, and models inside a single provider ecosystem, and that compatibility is the whole vulnerability.
A weaker sibling does the decoding: Inject an encrypted trace from a strong model into a weaker, less safeguarded model from the same provider and it decodes and emits the trace verbatim in plaintext. The capable model is never jailbroken directly. Recovered token counts match billed thinking tokens 1:1 on most queries.
Four attack vectors, not one: It circumvents anti-distillation across Anthropic, OpenAI, and Google. Decoding 315,320 blocks scraped from public repositories recovered 367 PII artifacts and 182 credentials. It exposes hazardous content the visible output refused, and it enables invisible prompt injections hidden entirely inside encrypted blocks to poison public agentic rollouts.
Why it matters: Teams publish session logs assuming the encrypted blobs are opaque, and they are not. The authors disclosed responsibly and propose cryptographic and system-level mitigations, but the immediate action is auditing what your published traces actually contain.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack