Quiet-STaR
Free while signed in. Answers cite the passages they came from.

Quiet-STaR generalizes the Self-Taught Reasoner (STaR) so that a language model learns to generate internal rationales between every token, not just for explicit QA problems.
Rationales per token: At each token position the model emits a latent "thought" that helps predict the next token; learnable boundary tokens delimit start/end of each thought.
Tokenwise parallel sampling: A specialized attention kernel generates all per-token thoughts in parallel, turning what would be a quadratic blowup into tractable training.
REINFORCE objective: Thoughts that improve next-token predictions are rewarded; the authors also use extended teacher forcing to stabilize training.
Zero-shot gains: On Mistral-7B, zero-shot GSM8K jumps from 5.9% to 10.9% and CommonsenseQA from 36.3% to 47.2% - improvements obtained purely from continued pretraining without any task-specific fine-tuning.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack