AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
Xiaomi's LLM-Core team reports MiMo-V2.6, an omni-modal MoE family (Pro at 1.02T total / 42B active, Flash at 310B / 15B active) trained by scaling RL compute across batch size, environments and grader compute.

Do LLMs Learn from Rewards in Context? : Rethinking the role of reward in In-Context Reinforcement Learning
Minchan Kwon, Junmo Kim and colleagues at KAIST test whether the reward in direct in-context reinforcement learning actually acts as a learning signal, and find that it barely does. NeurIPS 2026 Spotlight, Negative Results track.

Structuring MoE Expert Selection for Agentic Reinforcement Learning
Bolian Li (Apple and Purdue) with Ting-Yao Hu, Cheng-Yu Hsieh, Oncel Tuzel, Raviteja Vemulapalli and colleagues at Apple study how mixture-of-experts routing relates to agentic behavior and control it during RL post-training.

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
Xin Wang, Dayiheng Liu, Jianwei Zhang and colleagues at Alibaba Group (Qwen team, with Ohio State) present TRACE, an FP4 quantization framework for RL training of MoE language models that aligns training-side and rollout-side quantization.

ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning
Seil Kang, Hangoo Kang, Tarun Suresh and Azalia Mirhoseini (Stanford) with Yonsei, Korea University and Bespoke Labs present ThunderSyncRL, which streams gradient computation during agentic RL without introducing policy staleness.

Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps
Jiaxin Zhang, Chien-Sheng Wu and colleagues at Salesforce AI Research propose Prospective Hindsight (PH), a reweighting rule that up-weights rollouts where the agent's own prediction of the outcome disagreed with the verifier's verdict; the paper is accepted at NeurIPS 2026.

ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning
Kun Feng, Yuchen Fang and colleagues at ShanghaiTech University and Ant Group introduce ARISE, an agentic RL framework that turns rollout evidence into paired rubrics and skills, retires criteria once mastered, and samples tasks by estimated capability, raising Qwen3.5-27B from 23.4% to 45.6% on SkillsBench.

Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets
Shuze Daniel Liu (MIT), Claire Chen (Caltech), Jiuqi Wang (UVA), David Simchi-Levi (MIT) and Thorsten Joachims (Cornell) train a 30B LLM seller with RL to negotiate a catalog of substitutable products with several buyers at once under a shared turn budget, and report that it beats much larger frontier models on profit.

PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning
Jiaan Zhu, Wei Gao and colleagues at USTC and HKUST with Alibaba Group present PEARL, an asynchronous agentic RL system that combines elastic GPUs, temporary reuse of idle training GPUs, and per-workload choice between prefill-decode colocation and disaggregation to speed up multi-turn rollouts.

My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning
Yihua Zhu, Qianying Liu, Weixu Qiao and colleagues at Alibaba with Kyoto University propose FAULT, which converts an agent's own natural-language diagnosis of its errors into step-level credit that is anchored to the terminal reward in agentic RL.

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning
Dongwon Jung, Muhao Chen, Varun Chandrasekaran, Jaron Lanier and colleagues at UC Davis and Microsoft (with UW and Purdue) introduce ProVer, which gives step-level credit in agentic GRPO only to segments that a judge proposes and rollouts then confirm.

ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning
Zhenlong Dai, Jingyuan Chen and colleagues at Zhejiang University with Ant Group (NeurIPS 2026) introduce ToolSearcher, an RL framework for choosing tools from very large repositories through multi-turn search.

Towards Full Pipeline FP8 Reinforcement Learning for LLMs
Fanchao Chen and Shivaram Venkataraman (UW-Madison) with Haibin Lin and colleagues at ByteDance Seed show why RL training with FP8 in both rollout and training collapses, and fix it with Calibrated Clipping.

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning
Xincheng Yao (Shanghai Jiao Tong, intern at Tencent AI Platform) and colleagues propose GRAFT, which merges GRPO rollouts into a trajectory graph to estimate step-level advantages that follow the textbook definition.

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
Ruoqi Guo, Yi Liu and Leo Yu Zhang (Griffith University) with colleagues at NTU, UNSW, Deakin, George Mason and Wake Forest build RLCDAlignBench and test whether Jev, TypeSafe AI's model trained with reinforcement learning for calibrated decisions, can detect alignment failures with no task-specific training.

ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning
Ming Ma, Yi Zhu, Yiran Zhong, Steven Hoi and colleagues at Tongyi Lab (Alibaba), UCAS, UCLA and NTU propose ProCredit, which reruns a task's acceptance checks after every turn and rewards each turn by how much verified progress it made.

A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents
Aakash Kolekar, Sahika Genc and colleagues at Amazon Advertising and AWS Agentic AI study when to use SFT, RL or both for long-horizon advertising analytics agents, and turn the answer into a per-feature routing diagnostic (EMNLP 2026 Industry Track).

FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model
Jingxuan Xu, Gang Wu, Yanan Wu, Yutao Mou and colleagues (independent researchers with Peking, Nanjing and BUPT) propose FLARE, which trains a lightweight generative reward model to give step-level risk feedback to long-horizon coding agents during inference, SFT and RL.

Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
Yuan, Kang, Liu, Choi, Iyer, Jiang and Jaques (University of Washington and Stanford) propose MoDA, an online RL post-training method that counters alignment-induced mode collapse by conditioning one shared policy on numbered roles that are rewarded for producing outputs distinct from each other.

HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning
Hongliang Wei and colleagues at Harbin Institute of Technology and Alibaba Cloud train one policy across several agent harnesses and introduce HarnessBandit, an online scheduler that picks which harness to train on at each optimizer step.

EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning
Weiyuan Li and colleagues (Fudan University) propose EvoRS, where an agentic designer rewrites the reward system during RL training from on-policy rollouts and reward traces, instead of keeping rubrics fixed.

CovR: Coverage-Aware Hardware Verification via Reasoning-Guided Reinforcement Learning
Abdelatty, Nouh and Reda (Brown University) build CovR, an agentic testbench-generation system for RTL hardware verification that optimizes for coverage rather than functional correctness alone, and distill the resulting behavior into a student model with simulation-derived rewards.

EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning
Khomich, Hermansson and Hakimi treat tree-structured rollout construction in LLM RL as a compute-allocation problem, deriving where to branch from a law-of-total-variance decomposition of the local policy gradient rather than from policy entropy.

F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows
Bojian Xiong and a 14-author team score a DeepSearch run across its whole pipeline rather than only its final answer, and release a benchmark for reward models in that setting.