AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Yan Yu and colleagues find that a privileged teacher is not always reliable and that teacher supervision helps only at certain training stages, and propose RetireOPD, where the student drops the teacher on its own.

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning
Wenjie Liao, Liangjie Zhao and Zehong Cao train task generation, execution and evaluation jointly in UnifiedPlayers, rather than pairing self-generated trajectories with a static verifier.

Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization
Yingxuan Zhuang and colleagues separate two optimization axes in agent RL, how feedback is exploited within a trajectory and how trajectories are aggregated across a batch, and address each with BATON.

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
Juzheng Zhang and colleagues at AWS AI Labs and the University of Maryland introduce ActObs, which applies SFT loss to the environment observation tokens already present in agent trajectories rather than only to agent-authored action tokens.

Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA
Kailong Fan, Yichen Wu and colleagues (Harvard Medical School/MGH) show why majority-vote test-time RL collapses on medical multiple-choice QA and propose PROSE, which rewards reasoning steps instead of answer agreement.

Learning to Solve Hard Problems in RL for LLMs by Never Giving Up
Michael Noukhovitch, Hamish Ivison, Nathan Lambert and Aaron Courville (Mila and Ai2) show that RL for LLMs improves easy problems far more than hard ones, a pattern they call the Matthew Effect, and propose Never Give Up (NGU), which keeps sampling a problem until one rollout is correct.

Expert-Space Exploration in MoE Reinforcement Learning
Hongyi He and colleagues at Microsoft Research use the expert-routing choices of MoE models as a source of exploration during RL, and propose ESRL to perturb routing without degrading rollouts.

Beyond Solver Verdicts: Generative Reward Models for Autoformalization
Vikash Singh, Debargha Ganguly and colleagues (Case Western Reserve University with Amazon Web Services) show that solver verdicts cannot detect formal translations that are wrong but still return the expected verdict, and train a generative verifier that scores equivalence to the reference formalization.

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
Junyao Yang, Yucheng Shi and colleagues at Tencent's Hy Foundation Model Frontier team (with NUS, Georgia, Indiana and Maryland) train T1, a 122B-total MoE model, with RL in a real cloud shell for up to 300+ tool-call turns per task, and raise Terminal-Bench 2.1 from 43.8% to 64.0%.

Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning
Bin Lei and colleagues at Salesforce AI Research and the University of Minnesota place the forks of tree-structured RLVR rollouts at the step where the model's answer belief shifts most, which gives step-level credit where the outcome is still undecided.

Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection
Haoyue Liu, Xiaoyu Ma and Ye Chen (CUHK Shenzhen and FNii-Shenzhen) show that GRPO is structurally mismatched when the tool-subset action space is small enough to enumerate, and replace sampling with exact expectation over the full space.

Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward
Eshwar Reddy M (Testsigma) and Sourav Karmakar (Intuit India) argue that the constraint on reasoning RL outside formal domains is the absence of a scalable sound reward, derive the exchange rate between verifier quality and test-time compute, and measure it against executable ground truth.

TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards
Rui Sun, Zhan Shi and Bing He (independent researchers) train diagnostic reasoning agents by sampling an intervention, injecting it into a simulator, and generating the observations it would produce, so the hidden intervention supplies an oracle label for a task where real ground truth would require expert investigation.

Long-Horizon Language Model Reinforcement Learning via Progressive Point Matching
Preston Fu, Kevin Frans, Oleh Rybkin and Sergey Levine at UC Berkeley with Aviral Kumar at CMU give an unbiased dense-reward formulation, progressive point matching, that rewards partial progress at the segment level and scales exponentially better than sparse outcome rewards on long trajectories.

Elastic Horizon: Discovering the Effective Interaction Frontier in Agentic Reinforcement Learning
Gangyi Zhang in the Qwen Business Unit of Alibaba with USTC collaborators propose the effective interaction frontier hypothesis and Elastic Horizon, a closed-loop controller that sets an agent's interaction budget from the 90th percentile of successful trajectory lengths instead of a hand-set maximum.

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models
Wonje Jeung and colleagues at Yonsei University, with Carnegie Mellon, show that vision-language models used as reward functions for robot learning give different rewards to the same trajectory when the goal instruction is paraphrased, and release a benchmark that measures it.

PaperGym: Rubric-Centered Evolution for Research-Plan Generation
Yuhan Wang and colleagues at Zhejiang University, with Kaitao Song at Apple, turn each scientific paper into a full RL environment by synthesizing the question from goal and background while deriving the criteria from method and experiments, cutting criterion leakage to 3.7%.

DE-Venus: A Data-Efficient RLVR Framework for Large Language Models
Shenzhi Yang and colleagues at Zhejiang University with Ant Group, HKBU, NTU and Southeast University present DE-Venus, a framework that treats RLVR supervision as evolving state across data preparation and policy optimization, so that sample selection, weak supervision and label correction can be compared inside one system rather than as separate papers.

TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents
Jinwei Gan at Nanjing University introduces TIGPO, which keeps a persistent per-task transition graph across policy updates so that credit assignment for long-horizon agents can draw on transitions discovered by earlier policy versions rather than only the current batch.

Spurious Advantage Hidden in GRPO
Jiamian Wang and colleagues identify spurious advantage in GRPO, where a rollout that lands on the right answer by guessing receives the same high advantage magnitude as one that reasoned its way there, and propose SIGNBALANCE to remove it.

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
Zixun Huang, Kishan Panaganti, Haitao Mi and Leowei Liang propose FlowBalance, which lets a reasoning model learn from its own dense self-guidance but calibrates every guidance signal against the verifier's group advantage so false confidence gets reversed rather than reinforced.

Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs
Jiacheng Xu and colleagues at Nanyang Technological University with Skywork AI frame automatic test-case generation as an adversarial RL problem, where the generator must produce counterexamples targeted at the solver's current failure modes.

Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO
Hyun Bin Park and Du-Seong Chang isolate replay in GRPO down to a single primitive with two decisions, Headroom for what is still worth learning from and Drift for what is still compatible with the current policy.

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards
Leqi Zheng and colleagues propose Gradient-Aligned Reward, which builds a dense reasoning-aware reward by comparing each rollout's gradient direction to an expert-anchor gradient, using expert solutions already sitting in the training corpus.