🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning

Demystifying Long Chain-of-Thought Reasoning in LLMs

First page
Demystifying Long Chain-of-Thought Reasoning in LLMs
Paper summary

This work investigates how LLMs develop extended CoT reasoning, focusing on RL and compute scaling. Key insights include:

Ask this paper

Key points
01

Supervised fine-tuning (SFT) boosts performance – While not strictly necessary, SFT simplifies training and increases efficiency. Models fine-tuned with long CoT data achieve higher accuracy than those using short CoT sequences.

02

Reward shaping is crucial for stable RL – The study finds that naive RL approaches don’t always extend CoT length effectively. To address this, the authors introduce a cosine length-scaling reward with repetition penalties, which balances reasoning depth and prevents meaningless length increases.

03

Scaling verifiable reward signals – RL models trained with noisy, web-extracted “silver” supervision signals can generalize better to OOD tasks, such as STEM reasoning. Filtering such data is crucial to maintaining training stability.

04

Emergent reasoning abilities in base models – Skills like error correction and backtracking exist in base models but require careful RL incentives to be effectively utilized in complex tasks. This paper provides a structured roadmap for researchers looking to refine CoT training strategies for LLMs, highlighting how RL and reward tuning impact reasoning depth.

Every Monday
Get next week’s papers.
Subscribe on Substack