🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning

Quiet-STaR

First page
Quiet-STaR
Paper summary

Quiet-STaR generalizes the Self-Taught Reasoner (STaR) so that a language model learns to generate internal rationales between every token, not just for explicit QA problems.

Ask this paper

Key points
01

Rationales per token: At each token position the model emits a latent "thought" that helps predict the next token; learnable boundary tokens delimit start/end of each thought.

02

Tokenwise parallel sampling: A specialized attention kernel generates all per-token thoughts in parallel, turning what would be a quadratic blowup into tractable training.

03

REINFORCE objective: Thoughts that improve next-token predictions are rewarded; the authors also use extended teacher forcing to stabilize training.

04

Zero-shot gains: On Mistral-7B, zero-shot GSM8K jumps from 5.9% to 10.9% and CommonsenseQA from 36.3% to 47.2% - improvements obtained purely from continued pretraining without any task-specific fine-tuning.

Every Monday
Get next week’s papers.
Subscribe on Substack