Sharpening Tax in Post-Training

Changdae Oh, Qi Zeng, Qi Qi and colleagues at Meta Superintelligence Labs, with Sharon Li (UW-Madison) and Azalia Mirhoseini (Stanford), test whether RL post-training narrows what agentic models can solve, and find that base models with a light harness often cover more tasks than their post-trained versions when given enough samples.
Ask this paper
Finding. Pre-trained base models with a minimal inference harness work as tool-calling agents. Their pass@1 is far lower, but at large K their pass@K often exceeds that of the post-trained checkpoint, on BFCL v4 multi-turn, ACEBench and WebShop.
Mechanism. Post-training pushes each task toward always solved or never solved. This raises sampling efficiency and consistency and lowers the number of distinct tasks reachable under repeated sampling.
Sharpening Tax metric. A single number for the loss in test-time scalability after post-training. Across 14 base and post-trained pairs from four families (3B to 35B) and three benchmarks, 42 cases, the tax appears in most settings, can be estimated from a few rollouts and grows with model scale.
PTGS fix. Posterior-tempered group sampling sets the sampling temperature per prompt from its estimated difficulty. Used during RL in two agentic environments, it pays a smaller tax than fixed temperature and also improves pass@1.
Implication. The authors argue post-training reports should include a coverage diagnostic next to accuracy, because the agentic gains from RL mostly change how reliably tasks are solved rather than which tasks can be solved.
Abstract
An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.