Does RL Incentivize Reasoning in LLMs Beyond the Base Model?
Free while signed in. Answers cite the passages they came from.

This paper revisits a key assumption in recent LLM development: that Reinforcement Learning with Verifiable Rewards (RLVR) helps models acquire genuinely new reasoning capabilities. By analyzing models across tasks (math, code, vision) using pass@k metrics (with large k), the authors find that RLVR improves sample efficiency but does not expand reasoning capacity beyond the base model.
Key insight: RLVR-trained models do better at low *k* (e.g., pass@1), but as *k* increases (up to 256 or more), base models eventually match or outperform them. This suggests RLVR doesnât generate fundamentally new reasoning paths but just increases the likelihood of sampling already-existing correct ones.
Reasoning already in the base: RLVR models' successful CoTs are shown to be present within the base model's sampling distribution. Perplexity analyses confirm that RL outputs are often high-probability continuations for the base model.
Efficiency vs. exploration: RLVR narrows the modelâs exploration space, improving efficiency but shrinking its coverage of diverse reasoning paths, thereby reducing overall problem-solving reach at scale.
Distillation helps more: Unlike RLVR, distillation from a stronger teacher model (e.g., DeepSeek-R1) introduces genuinely new reasoning patterns, expanding the modelâs capabilities.
Algorithmic limits: Across PPO, GRPO, Reinforce++, etc., RL algorithms offer similar sample-efficiency improvements, but none closes the gap to the base modelâs pass@256âhighlighting the limits of current RL strategies.
Get next weekâs papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack