🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
242 papers · Reinforcement LearningClear filters →
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

Yan Yu and colleagues find that a privileged teacher is not always reliable and that teacher supervision helps only at certain training stages, and propose RetireOPD, where the student drops the teacher on its own.

25Training
UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

Wenjie Liao, Liangjie Zhao and Zehong Cao train task generation, execution and evaluation jointly in UnifiedPlayers, rather than pairing self-generated trajectories with a static verifier.

26Reasoning
Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization

Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization

Yingxuan Zhuang and colleagues separate two optimization axes in agent RL, how feedback is exploited within a trajectory and how trajectories are aggregated across a batch, and address each with BATON.

27Agents
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

Juzheng Zhang and colleagues at AWS AI Labs and the University of Maryland introduce ActObs, which applies SFT loss to the environment observation tokens already present in agent trajectories rather than only to agent-authored action tokens.

28Reinforcement Learning
Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA

Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA

Kailong Fan, Yichen Wu and colleagues (Harvard Medical School/MGH) show why majority-vote test-time RL collapses on medical multiple-choice QA and propose PROSE, which rewards reasoning steps instead of answer agreement.

29Reinforcement Learning
Learning to Solve Hard Problems in RL for LLMs by Never Giving Up

Learning to Solve Hard Problems in RL for LLMs by Never Giving Up

Michael Noukhovitch, Hamish Ivison, Nathan Lambert and Aaron Courville (Mila and Ai2) show that RL for LLMs improves easy problems far more than hard ones, a pattern they call the Matthew Effect, and propose Never Give Up (NGU), which keeps sampling a problem until one rollout is correct.

30Reinforcement Learning
Expert-Space Exploration in MoE Reinforcement Learning

Expert-Space Exploration in MoE Reinforcement Learning

Hongyi He and colleagues at Microsoft Research use the expert-routing choices of MoE models as a source of exploration during RL, and propose ESRL to perturb routing without degrading rollouts.

31Architecture
Beyond Solver Verdicts: Generative Reward Models for Autoformalization

Beyond Solver Verdicts: Generative Reward Models for Autoformalization

Vikash Singh, Debargha Ganguly and colleagues (Case Western Reserve University with Amazon Web Services) show that solver verdicts cannot detect formal translations that are wrong but still return the expected verdict, and train a generative verifier that scores equivalence to the reference formalization.

32Reinforcement Learning
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

Junyao Yang, Yucheng Shi and colleagues at Tencent's Hy Foundation Model Frontier team (with NUS, Georgia, Indiana and Maryland) train T1, a 122B-total MoE model, with RL in a real cloud shell for up to 300+ tool-call turns per task, and raise Terminal-Bench 2.1 from 43.8% to 64.0%.

33Reinforcement Learning
Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning

Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning

Bin Lei and colleagues at Salesforce AI Research and the University of Minnesota place the forks of tree-structured RLVR rollouts at the step where the model's answer belief shifts most, which gives step-level credit where the outcome is still undecided.

34Reinforcement Learning
Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection

Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection

Haoyue Liu, Xiaoyu Ma and Ye Chen (CUHK Shenzhen and FNii-Shenzhen) show that GRPO is structurally mismatched when the tool-subset action space is small enough to enumerate, and replace sampling with exact expectation over the full space.

35Reinforcement Learning
Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward

Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward

Eshwar Reddy M (Testsigma) and Sourav Karmakar (Intuit India) argue that the constraint on reasoning RL outside formal domains is the absence of a scalable sound reward, derive the exchange rate between verifier quality and test-time compute, and measure it against executable ground truth.

36Reinforcement Learning
TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards

TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards

Rui Sun, Zhan Shi and Bing He (independent researchers) train diagnostic reasoning agents by sampling an intervention, injecting it into a simulator, and generating the observations it would produce, so the hidden intervention supplies an oracle label for a task where real ground truth would require expert investigation.

37Reinforcement Learning
Long-Horizon Language Model Reinforcement Learning via Progressive Point Matching

Long-Horizon Language Model Reinforcement Learning via Progressive Point Matching

Preston Fu, Kevin Frans, Oleh Rybkin and Sergey Levine at UC Berkeley with Aviral Kumar at CMU give an unbiased dense-reward formulation, progressive point matching, that rewards partial progress at the segment level and scales exponentially better than sparse outcome rewards on long trajectories.

38Reinforcement Learning
Elastic Horizon: Discovering the Effective Interaction Frontier in Agentic Reinforcement Learning

Elastic Horizon: Discovering the Effective Interaction Frontier in Agentic Reinforcement Learning

Gangyi Zhang in the Qwen Business Unit of Alibaba with USTC collaborators propose the effective interaction frontier hypothesis and Elastic Horizon, a closed-loop controller that sets an agent's interaction budget from the 90th percentile of successful trajectory lengths instead of a hand-set maximum.

39Agents
Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

Wonje Jeung and colleagues at Yonsei University, with Carnegie Mellon, show that vision-language models used as reward functions for robot learning give different rewards to the same trajectory when the goal instruction is paraphrased, and release a benchmark that measures it.

40Multimodal
PaperGym: Rubric-Centered Evolution for Research-Plan Generation

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

Yuhan Wang and colleagues at Zhejiang University, with Kaitao Song at Apple, turn each scientific paper into a full RL environment by synthesizing the question from goal and background while deriving the criteria from method and experiments, cutting criterion leakage to 3.7%.

41Reinforcement Learning
DE-Venus: A Data-Efficient RLVR Framework for Large Language Models

DE-Venus: A Data-Efficient RLVR Framework for Large Language Models

Shenzhi Yang and colleagues at Zhejiang University with Ant Group, HKBU, NTU and Southeast University present DE-Venus, a framework that treats RLVR supervision as evolving state across data preparation and policy optimization, so that sample selection, weak supervision and label correction can be compared inside one system rather than as separate papers.

42Data
TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents

TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents

Jinwei Gan at Nanjing University introduces TIGPO, which keeps a persistent per-task transition graph across policy updates so that credit assignment for long-horizon agents can draw on transitions discovered by earlier policy versions rather than only the current batch.

43Agents
Spurious Advantage Hidden in GRPO

Spurious Advantage Hidden in GRPO

Jiamian Wang and colleagues identify spurious advantage in GRPO, where a rollout that lands on the right answer by guessing receives the same high advantage magnitude as one that reasoned its way there, and propose SIGNBALANCE to remove it.

44Reinforcement Learning
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

Zixun Huang, Kishan Panaganti, Haitao Mi and Leowei Liang propose FlowBalance, which lets a reasoning model learn from its own dense self-guidance but calibrates every guidance signal against the verifier's group advantage so false confidence gets reversed rather than reinforced.

45Reinforcement Learning
Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs

Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs

Jiacheng Xu and colleagues at Nanyang Technological University with Skywork AI frame automatic test-case generation as an adversarial RL problem, where the generator must produce counterexamples targeted at the solver's current failure modes.

46Reinforcement Learning
Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO

Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO

Hyun Bin Park and Du-Seong Chang isolate replay in GRPO down to a single primitive with two decisions, Headroom for what is still worth learning from and Drift for what is still compatible with the current policy.

47Reinforcement Learning
Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

Leqi Zheng and colleagues propose Gradient-Aligned Reward, which builds a dense reasoning-aware reward by comparing each rollout's gradient direction to an expert-anchor gradient, using expert solutions already sitting in the training corpus.

48Safety
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026