🚀NEW LABGetting Started with Claude AgentsStart lab
Training · Reasoning · Reinforcement Learning

Thinking Mid-training: RL of Interleaved Reasoning

Paper preview
Thinking Mid-training: RL of Interleaved Reasoning
Paper summary

Meta FAIR addresses the gap between pretraining (no explicit reasoning) and post-training (reasoning-heavy) with an intermediate SFT+RL mid-training phase. The approach annotates pretraining data with interleaved reasoning traces, then uses supervised fine-tuning followed by RL to teach models when and how to think during continued pretraining. Applied to Llama-3-8B, the full pipeline achieves a 3.2x improvement on reasoning benchmarks compared to direct RL post-training, demonstrating that reasoning benefits from being trained as native behavior early in the pipeline.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack