🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
242 papers · Reinforcement LearningClear filters →
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Xiaomi's LLM-Core team reports MiMo-V2.6, an omni-modal MoE family (Pro at 1.02T total / 42B active, Flash at 310B / 15B active) trained by scaling RL compute across batch size, environments and grader compute.

01Reinforcement Learning
Do LLMs Learn from Rewards in Context? : Rethinking the role of reward in In-Context Reinforcement Learning

Do LLMs Learn from Rewards in Context? : Rethinking the role of reward in In-Context Reinforcement Learning

Minchan Kwon, Junmo Kim and colleagues at KAIST test whether the reward in direct in-context reinforcement learning actually acts as a learning signal, and find that it barely does. NeurIPS 2026 Spotlight, Negative Results track.

02Reinforcement Learning
Structuring MoE Expert Selection for Agentic Reinforcement Learning

Structuring MoE Expert Selection for Agentic Reinforcement Learning

Bolian Li (Apple and Purdue) with Ting-Yao Hu, Cheng-Yu Hsieh, Oncel Tuzel, Raviteja Vemulapalli and colleagues at Apple study how mixture-of-experts routing relates to agentic behavior and control it during RL post-training.

03Architecture
TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

Xin Wang, Dayiheng Liu, Jianwei Zhang and colleagues at Alibaba Group (Qwen team, with Ohio State) present TRACE, an FP4 quantization framework for RL training of MoE language models that aligns training-side and rollout-side quantization.

04Efficiency
ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning

ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning

Seil Kang, Hangoo Kang, Tarun Suresh and Azalia Mirhoseini (Stanford) with Yonsei, Korea University and Bespoke Labs present ThunderSyncRL, which streams gradient computation during agentic RL without introducing policy staleness.

05Agents
Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps

Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps

Jiaxin Zhang, Chien-Sheng Wu and colleagues at Salesforce AI Research propose Prospective Hindsight (PH), a reweighting rule that up-weights rollouts where the agent's own prediction of the outcome disagreed with the verifier's verdict; the paper is accepted at NeurIPS 2026.

06Reinforcement Learning
ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning

ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning

Kun Feng, Yuchen Fang and colleagues at ShanghaiTech University and Ant Group introduce ARISE, an agentic RL framework that turns rollout evidence into paired rubrics and skills, retires criteria once mastered, and samples tasks by estimated capability, raising Qwen3.5-27B from 23.4% to 45.6% on SkillsBench.

07Agents
Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets

Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets

Shuze Daniel Liu (MIT), Claire Chen (Caltech), Jiuqi Wang (UVA), David Simchi-Levi (MIT) and Thorsten Joachims (Cornell) train a 30B LLM seller with RL to negotiate a catalog of substitutable products with several buyers at once under a shared turn budget, and report that it beats much larger frontier models on profit.

08Reinforcement Learning
PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning

PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning

Jiaan Zhu, Wei Gao and colleagues at USTC and HKUST with Alibaba Group present PEARL, an asynchronous agentic RL system that combines elastic GPUs, temporary reuse of idle training GPUs, and per-workload choice between prefill-decode colocation and disaggregation to speed up multi-turn rollouts.

09Agents
My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning

My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning

Yihua Zhu, Qianying Liu, Weixu Qiao and colleagues at Alibaba with Kyoto University propose FAULT, which converts an agent's own natural-language diagnosis of its errors into step-level credit that is anchored to the terminal reward in agentic RL.

10Agents
Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

Dongwon Jung, Muhao Chen, Varun Chandrasekaran, Jaron Lanier and colleagues at UC Davis and Microsoft (with UW and Purdue) introduce ProVer, which gives step-level credit in agentic GRPO only to segments that a judge proposes and rollouts then confirm.

11Agents
ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

Zhenlong Dai, Jingyuan Chen and colleagues at Zhejiang University with Ant Group (NeurIPS 2026) introduce ToolSearcher, an RL framework for choosing tools from very large repositories through multi-turn search.

12Reinforcement Learning
Towards Full Pipeline FP8 Reinforcement Learning for LLMs

Towards Full Pipeline FP8 Reinforcement Learning for LLMs

Fanchao Chen and Shivaram Venkataraman (UW-Madison) with Haibin Lin and colleagues at ByteDance Seed show why RL training with FP8 in both rollout and training collapses, and fix it with Calibrated Clipping.

13Reinforcement Learning
Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

Xincheng Yao (Shanghai Jiao Tong, intern at Tencent AI Platform) and colleagues propose GRAFT, which merges GRPO rollouts into a trajectory graph to estimate step-level advantages that follow the textbook definition.

14Agents
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

Ruoqi Guo, Yi Liu and Leo Yu Zhang (Griffith University) with colleagues at NTU, UNSW, Deakin, George Mason and Wake Forest build RLCDAlignBench and test whether Jev, TypeSafe AI's model trained with reinforcement learning for calibrated decisions, can detect alignment failures with no task-specific training.

15Safety
ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning

ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning

Ming Ma, Yi Zhu, Yiran Zhong, Steven Hoi and colleagues at Tongyi Lab (Alibaba), UCAS, UCLA and NTU propose ProCredit, which reruns a task's acceptance checks after every turn and rewards each turn by how much verified progress it made.

16Reinforcement Learning
A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents

A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents

Aakash Kolekar, Sahika Genc and colleagues at Amazon Advertising and AWS Agentic AI study when to use SFT, RL or both for long-horizon advertising analytics agents, and turn the answer into a per-feature routing diagnostic (EMNLP 2026 Industry Track).

17Agents
FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model

FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model

Jingxuan Xu, Gang Wu, Yanan Wu, Yutao Mou and colleagues (independent researchers with Peking, Nanjing and BUPT) propose FLARE, which trains a lightweight generative reward model to give step-level risk feedback to long-horizon coding agents during inference, SFT and RL.

18Reinforcement Learning
Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning

Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning

Yuan, Kang, Liu, Choi, Iyer, Jiang and Jaques (University of Washington and Stanford) propose MoDA, an online RL post-training method that counters alignment-induced mode collapse by conditioning one shared policy on numbered roles that are rewarded for producing outputs distinct from each other.

19Safety
HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning

HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning

Hongliang Wei and colleagues at Harbin Institute of Technology and Alibaba Cloud train one policy across several agent harnesses and introduce HarnessBandit, an online scheduler that picks which harness to train on at each optimizer step.

20Agents
EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning

EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning

Weiyuan Li and colleagues (Fudan University) propose EvoRS, where an agentic designer rewrites the reward system during RL training from on-policy rollouts and reward traces, instead of keeping rubrics fixed.

21Reinforcement Learning
CovR: Coverage-Aware Hardware Verification via Reasoning-Guided Reinforcement Learning

CovR: Coverage-Aware Hardware Verification via Reasoning-Guided Reinforcement Learning

Abdelatty, Nouh and Reda (Brown University) build CovR, an agentic testbench-generation system for RTL hardware verification that optimizes for coverage rather than functional correctness alone, and distill the resulting behavior into a student model with simulation-derived rewards.

22Reasoning
EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning

EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning

Khomich, Hermansson and Hakimi treat tree-structured rollout construction in LLM RL as a compute-allocation problem, deriving where to branch from a law-of-total-variance decomposition of the local policy gradient rather than from policy entropy.

23Reinforcement Learning
F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows

F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows

Bojian Xiong and a 14-author team score a DeepSearch run across its whole pipeline rather than only its final answer, and release a benchmark for reward models in that setting.

24Evaluation
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026