🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Safety · Reasoning · Evaluation

Invisible Reasoning

Free while signed in. Answers cite the passages they came from.

First page
Invisible Reasoning
The curator’s take

Chain-of-thought monitoring rests on the assumption that a model expresses its reasoning in its output tokens. This work demonstrates a concrete failure of that assumption in models shipping today.

Key points
01

Filler tokens carry computation: Across 13 frontier models and three tasks, many models improve significantly when given semantically irrelevant filler tokens, with accuracy gains of up to 13 percentage points, and the benefit depends on which tokens are used.

02

Hidden objectives are reachable: Filler tokens let Claude Opus 4.5 satisfy a hidden modular arithmetic constraint without sacrificing accuracy on its primary task, showing that invisible reasoning can serve goals a CoT monitor never sees.

03

Training does not induce it: Reinforcement learning gives Qwen3-235B strong preferences over filler token content, but neither RL nor supervised fine-tuning produces a filler token benefit that persists at test time, so this is a property of the frontier models rather than a trick you can bolt on.

04

Why it matters: If consequential computation already happens with no interpretable trace in the output tokens, then CoT-based oversight is a partial signal, and safety cases built on reading the reasoning need to account for what is not written down.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack