🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning

Chain-of-Thought Reasoning Without Prompting

First page
Chain-of-Thought Reasoning Without Prompting
Paper summary

DeepMind shows that LLMs often *already* emit chain-of-thought reasoning in alternative decoding paths, and that selecting those paths via confidence lifts reasoning accuracy with no prompt engineering.

Ask this paper

Key points
01

Alternative decoding: Instead of taking the top-1 greedy token, the method considers top-k alternative first tokens and runs the full decode from each, exposing paths that naturally produce CoT.

02

Confidence as a signal: When a decoded path contains genuine step-by-step reasoning, the model's final-answer confidence is noticeably higher - this acts as an automatic selector for the CoT path.

03

Benchmark gains: The technique substantially outperforms standard greedy decoding across arithmetic, commonsense, and symbolic reasoning benchmarks without any prompt modification.

04

Reframing: Reasoning is positioned as already latent in pretrained LLMs - prompt engineering is one way to surface it, but decoding-time selection is an equally valid (and complementary) lever.

Every Monday
Get next week’s papers.
Subscribe on Substack