🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 25, 2026
Reasoning

Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning

First page
Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning
The curator’s take

Ismail Labiad, Rémi Munos, Julia Kempe and colleagues at Meta FAIR with Université Paris-Saclay and NYU train a small concept generator with RL so that the hints it writes raise the pass@k of a larger frozen answer model.

Ask this paper

Key points
01

Prior gains do not reproduce. Re-running GuidedSampling on all 25 Qwen2.5 size pairs on MATH500 with an exploratory repeated-sampling baseline removes its reported advantage; the original baseline was sampled with restrictive settings.

02

Single-trajectory concepts. Emitting many problem-specific concepts in one generation, then splitting the answer budget across them, gives consistent pass@50 gains on hard MATH subsets without any training.

03

RL over concepts. A Qwen2.5-7B concept generator rewarded by the downstream success of a frozen Qwen2.5-32B doubles pass@128 on 1k held-out hard DeepMath problems, from 19.0% to 39.2% at equal answer-generation compute.

04

Small beats large. The trained 7B generator beats the untuned 32B generator (39.2% vs 33.8%), and transfers without retraining to Llama-3.3-70B (26.2% to 34.3% pass@128).

05

Relevance accounts for most of it. Concepts from an unrelated problem reach 23.9%, so problem-specific content produces about 15 of the 20-point gain.

Abstract

Large language models increasingly tackle hard reasoning problems by spending more test-time compute, yet the dominant strategy remains naive repeated sampling: draw many independent solutions and hope one is correct. Because such sampling explores only through local decoding noise, it tends to produce many near duplicate attempts rather than genuinely different ideas. We ask whether exploration can instead be steered at a semantic level, by first sampling problem specific concepts, hints, or strategies and then conditioning answer generation on them. We refine this into a simple, more exploratory procedure that emits many diverse concepts in a single trajectory, and evaluate it on hard problems where repeated sampling struggles. We then go a step further and make concept generation trainable: a small concept generator is optimized with reinforcement learning so that its concepts maximize the downstream success of a larger, frozen answer generator. On hard mathematical reasoning problems, the trained concept generator substantially improves the answer generator's pass@k over naive repeated sampling at the same answer generation allocation, surpasses concepts drawn from much larger untuned models, and transfers to answer generators it was never trained against, including a model from a different family. A small model can thus be trained into an effective, reusable search policy for a much larger one.

Every Monday
Get next week’s papers.
Subscribe on Substack