Chain-of-Thought Reasoning Without Prompting

DeepMind shows that LLMs often *already* emit chain-of-thought reasoning in alternative decoding paths, and that selecting those paths via confidence lifts reasoning accuracy with no prompt engineering.
Ask this paper
Alternative decoding: Instead of taking the top-1 greedy token, the method considers top-k alternative first tokens and runs the full decode from each, exposing paths that naturally produce CoT.
Confidence as a signal: When a decoded path contains genuine step-by-step reasoning, the model's final-answer confidence is noticeably higher - this acts as an automatic selector for the CoT path.
Benchmark gains: The technique substantially outperforms standard greedy decoding across arithmetic, commonsense, and symbolic reasoning benchmarks without any prompt modification.
Reframing: Reasoning is positioned as already latent in pretrained LLMs - prompt engineering is one way to surface it, but decoding-time selection is an equally valid (and complementary) lever.