🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 6, 2026
Reinforcement Learning · Safety · Agents

Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps

First page
Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps
The curator’s take

Jiaxin Zhang, Chien-Sheng Wu and colleagues at Salesforce AI Research propose Prospective Hindsight (PH), a reweighting rule that up-weights rollouts where the agent's own prediction of the outcome disagreed with the verifier's verdict; the paper is accepted at NeurIPS 2026.

Ask this paper

Key points
01

Hindsight trap. Retrospective RL assigns credit only after the outcome is known, so overconfident failures and underconfident successes are treated the same as calibrated rollouts with the same label.

02

Method. Before scoring, the same policy is asked with a different prompt to predict the outcome. The gap between prediction and verifier score sets a stop-gradient weight on the base loss, which works with GRPO, on-policy distillation or both without changing rollout sampling.

03

Self-extinguishing curriculum. Because the predictor shares weights with the policy, calibration improves as a side effect of training and the extra weighting shrinks as the prediction gap closes, without a separate calibration loss or schedule.

04

Results. PH improves task performance and calibration on single-turn verifiable tasks and on a multi-turn personal-agent task with a GPT-4.1 user simulator, with the largest gain on on-policy distillation, and the gains hold on OLMo-3-7B-Instruct and gpt-oss-20B.

Abstract

Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradient. We introduce Prospective Hindsight (PH), a self-calibrating training principle that augments any retrospective base method with a signal derived from the gap between the agent's prospective prediction (before feedback) and the retrospective evaluation (after feedback). This per-rollout surprise identifies samples where the agent's self-model is most inaccurate and amplifies their gradient contribution through a stop-gradient surprise-weighted advantage. Since the prospective predictor shares parameters with the policy, the two co-evolve, progressively shifting focus to the agent's remaining blind spots. We connect this principle to a privileged-information gap and show that minimizing the surprise residual provides a descent pathway on the agent's miscalibration rate; calibration thus emerges as a byproduct of optimization rather than from an added objective. On single-turn verifiable tasks and a multi-turn personal-agent task (under GRPO, on-policy distillation, and their combination), PH improves both task performance and calibration, with consistent gains across model scales. Notably, the dominant miscalibration mode shifts structurally between regimes, overconfident failures in single-turn, underconfident successes in multi-turn, yet the same training principle addresses both successfully.

Every Monday
Get next week’s papers.
Subscribe on Substack