🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 21, 2026
Reasoning

Unlocking the Unsolvable: Teacher-Guided Curriculum for Data-Efficient RLVR

First page
Unlocking the Unsolvable: Teacher-Guided Curriculum for Data-Efficient RLVR
The curator’s take

Zhu (independent) and Han (Amazon, work done independently) show that RLVR problems a model cannot yet solve, which normally yield zero learning signal, can be made useful by giving partial teacher reasoning traces and withdrawing them step by step.

Ask this paper

Key points
01

Backward chaining. Partial traces from a stronger model create graded difficulty; the curriculum shortens the provided prefix until the student solves the problem unaided.

02

Data efficiency. Training on only 128 unsolvable problems matches or beats GRPO on the full 2,000-problem corpus on a nine-benchmark average, about 16 times more data-efficient, for both base models tested.

03

Reasoning boundary. Pass@k at large k improves substantially, which indicates the method expands what the model can solve rather than only sharpening existing solutions.

04

Monotone Frontier Curriculum. The authors identify a distribution-shift cost specific to unsolvable-only training and propose MFC, which moves monotonically toward unguided solving and outperforms prior curriculum methods.

05

Venue. Accepted to Findings of EMNLP 2026.

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has shown remarkable success in improving the mathematical reasoning of large language models. Yet problems beyond the model's current capability, where rollouts uniformly fail and no learning signal is produced, are structurally wasted despite marking the most informative training frontier. We show that these otherwise-inert problems can be unlocked via teacher-guided curriculum learning: partial reasoning traces from a stronger model create a graded difficulty landscape, and a backward-chaining curriculum progressively withdraws guidance until the model solves problems unaided. Training on only 128 unsolvable problems matches or exceeds GRPO trained on a full 2,000-problem corpus (~16x data efficiency) on the nine-benchmark average for both base models, while substantially expanding the reasoning boundary measured by pass@k at large k. Furthermore, we identify a distribution-shift cost that is particularly acute in the unsolvable-only regime and propose Monotone Frontier Curriculum (MFC), a method that monotonically drives training toward unguided solving, consistently outperforming existing curriculum methods.

Every Monday
Get next week’s papers.
Subscribe on Substack