🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning

Training LLMs with Pause Tokens

First page
Training LLMs with Pause Tokens
Paper summary

CMU shows that adding a learnable `<pause>` token during both pretraining and fine-tuning gives the model extra "thinking time" and improves reasoning.

Ask this paper

Key points
01

Learnable pause token: Inserts a `<pause>` token into the input; the model processes these tokens but doesn't treat them as meaningful content, letting it compute more before answering.

02

CommonsenseQA and math gains: Produces measurable performance gains on CommonsenseQA and math word problems - both tasks that benefit from extra internal computation.

03

Pretraining is required: The benefit only materializes if pauses are introduced in both pretraining and fine-tuning - adding them only at inference doesn't work.

04

Compute-aware decoding: Positions pause tokens as a simple inference-time knob for trading compute against accuracy, foreshadowing many 2024 "thinking time" tricks.

Every Monday
Get next week’s papers.
Subscribe on Substack