🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
242 papers · Reinforcement LearningClear filters →
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

Shubham Gandhi (CMU) with Saurabh Goyal, Kiran Kate and Yara Rizk at IBM Research tackle the outcome-blind setting, where long-horizon agent tasks have no programmatic checker, by redistributing a once-per-trajectory rubric judgment over the steps that earned it.

49Reinforcement Learning
Cliff: Learning Process Rewards from the First Mistake

Cliff: Learning Process Rewards from the First Mistake

Peixuan Han, Runhui Wang and colleagues at AWS propose Cliff, a reward shaping method that asks an off-the-shelf teacher LLM to find only the first mistake in a rollout, then converts that single index into dense token-level advantages.

50Reinforcement Learning
Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment

Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment

Chenyu Zhou and colleagues make an uncomfortable argument for multi-turn agentic RL: given a terminal-state verifier, spreading reward uniformly beats every attempt to target the turns that mattered, and they name the quantity that predicts when this holds.

51Reinforcement Learning
Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

Liming Pu and colleagues at Alibaba Research argue that the widely assumed ceiling on outcome-only RL for small open agent models is a practice artifact rather than a property of the method, and present CANOPY, a stripped-down protocol that tops the AppWorld leaderboard with a 14B policy trained purely through environment interaction.

52Agents
ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

Fanrui Zhang and a large Alibaba-affiliated team propose ARISE-RL, a co-evolutionary loop in which a task and rubric Generator and a reasoning Solver train each other, replacing the verifiable gold answer that open-ended agentic RL does not have.

53Agents
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

Yunpeng Ba and colleagues (Huawei Noah's Ark Lab, City University of Hong Kong) explain when Evolution Strategies beat GRPO for LLM reasoning, tying the advantage to reasoning coverage rather than to raw reward.

54Reinforcement Learning
Reflexion: Language Agents with Verbal Reinforcement Learning

Reflexion: Language Agents with Verbal Reinforcement Learning

Take the real reward signal from the environment and write it back into the context as words. Reflexion is where a failed episode stops being wasted, which is the seed of everything in the self-improving half of this list.

55Reinforcement Learning
Recurrent Looped Transformer

Recurrent Looped Transformer

Yifan Zhang proposes the Recurrent Looped Transformer, in which a causal encoder builds global key-value memory and a recurrent decoder carries its final hidden state and sliding-window cache across every prompt and response token, so the depth of the computation path grows with sequence length while the number of blocks per token stays fixed. The report is a design specification and contains no experimental results.

56Architecture
Reinforcing Agents with Collective Skills

Reinforcing Agents with Collective Skills

Binfeng Xu, Yi Dong, Jan Kautz and colleagues at NVIDIA build Skill2Env, a pipeline that compiles public Agent Skills into executable RL environments for terminal agents, each with programmatic tests and a behavioral rubric drawn from the Skill's own quality criteria.

57Agents
Agent Lightning v1.0

Agent Lightning v1.0

Modern agents run inside a harness that owns tools, context, and control flow. When you want to train one, that ownership becomes the problem: the harness runs the environment loop and the trainer only ever sees LLM request and response pairs. This work from Microsoft treats that boundary as the integration point instead of an obstacle.

58Reinforcement Learning
ClawGym II

ClawGym II

If you want to train agents inside the harness they already run in, this is the black-box version of that idea. ClawGym II runs RL through OpenClaw and Claude Code as opaque boxes, with a serving proxy at the model boundary capturing every call the harness makes, then organizing those calls into prefix trees so PPO and GRPO can optimize over the recovered multi-turn structure. Qwen3-30A3B gains 9.98 points of Pass@1 through OpenClaw and 14.81 through Claude Code, stable across 200 to 400 optimization steps. Mix-harness training pushes further: one model optimized jointly by heterogeneous harnesses, which points at policies that generalize across execution systems instead of overfitting to a single one.

59Reinforcement Learning
Harness-R1

Harness-R1

Agents accumulate interaction trajectories during deployment and then leave them unused, because their behavior stays fixed. Those trajectories can improve the harness that constructs context, mediates tools, validates actions, and recovers execution, and this work makes that editing a learned capability.

60Agents
Molt

Molt

Agentic RL research is constant algorithm modification, and in mainstream frameworks each change threads through layers of trainer, distributed backend, and rollout glue. NVIDIA's Molt is a PyTorch-native training framework built to make that cost small.

61Reinforcement Learning
Role Drift

Role Drift

End-to-end RL improves the accuracy of a multi-module LLM pipeline without constraining how the modules divide labor internally. Harvard and MIT name the resulting failure mode, Role Drift, where a module preserves or improves end-task performance while abandoning its assigned role through shortcuts that system-level evaluation cannot see. Two instances showed up. A decomposer meant to split a question into sub-questions for a separate solver instead plants the answer inside them, and a reader meant to answer from retrieved passages instead falls back on parametric memory. Hold the decomposer to its role and 86% of the apparent RL gain disappears. Role Anchor, the proposed regularizer, preserves how the role prompt shifts a module's next-token predictions relative to a neutral prompt, and gradient analysis suggests it reduces alignment with the drift direction rather than simply suppressing learning.

62Reinforcement Learning
The Self-Speculating Agent

The Self-Speculating Agent

Agents spend a large share of wall-clock time waiting on tool results. Speculation hides that latency by predicting and pre-executing the next call, but external draft models and cached traces model a different policy, so they miss too often to help. UC Santa Barbara and LinkedIn identify this speculator-agent gap and unify both roles in one model. It runs in agent mode to solve the task and in speculator mode to predict its next tool call from a partial trajectory, fully reusing the prefix KV cache. Joint agent-speculator reinforcement learning derives speculation targets from the agent's own rollouts and alternates updates between the two modes. Next tool-call Hit@1 rises from 44.1 to 61.2 for Qwen3-4B and from 48.9 to 66.3 for Qwen3.5-4B, with agent task success preserved.

63Agents
GFlowRL

GFlowRL

Reward-maximizing RL tends to collapse large reasoning models onto a single dominant mode, and GFlowNet-style training is appealing because it matches reward distributions and keeps diverse reasoning paths. GFlowRL scales this to modern post-training by replacing the hard-to-learn partition function with an in-batch Monte Carlo estimate computed from the rollout group the pipeline already produces. It is the first GFlowNet-style RL algorithm to train stably across both dense and sparse architectures, reaching a 2048 Codeforces rating at 14B and outperforming prior methods on math, code, and adversarial red-teaming benchmarks like AdvBench and HarmBench.

64Reinforcement Learning
RLVR Meets Human Likeness

RLVR Meets Human Likeness

RL with verifiable rewards only optimizes what you can objectively score, so style, structure, and diversity quietly collapse and reward hacking creeps in. This MIT work adds an adversarial discriminator trained on human demonstrations as a learned proxy for the human output distribution, and the generator maximizes both task accuracy and that human-likeness signal. Across bug fixing, story generation, and a reward-hacking benchmark, it preserves RLVR's accuracy gains while restoring the fuzzy properties it usually destroys, with misbehavior nearly disappearing.

65Reinforcement Learning
Red Queen Gödel Machine

Red Queen Gödel Machine

Self-improving agents are only as strong as the evaluator scoring them, and most systems freeze that evaluator in place, so improvement stalls the moment the judge stops getting harder. The Red Queen Gödel Machine makes the evaluator part of the search itself, letting agents and the criteria that judge them co-evolve. --- ---

66Agents
The Verification Horizon

The Verification Horizon

Reinforcement learning for coding agents lives or dies on the reward signal, and this Qwen work argues there is no silver bullet. As policy capability grows, any fixed reward function eventually gets gamed, so verification has to co-evolve with the generator it scores. ---

67Reinforcement Learning
RLMF

RLMF

LLMs routinely hallucinate with high confidence, miss their own knowledge boundaries, and misreport uncertainty, and most fixes bolt calibration on from the outside. RLMF, a Google and Yale collaboration, instead turns the model’s own metacognition into the training signal. ---

68Reinforcement Learning
A Pinch of Human Data

A Pinch of Human Data

Self-play reinforcement learning can train driving policies with no human data at all, swapping expensive human demonstrations for cheap large-scale simulation. The catch is that pure self-play tends to discover effective but alien driving conventions that real people cannot work with, and the usual fixes lean on brittle reward engineering and domain randomization.

69Reinforcement Learning
From Trainee to Trainer

From Trainee to Trainer

Who should design the training environment for an RL agent, the practitioner or the policy itself? RL pipelines for LLMs usually rely on manually redesigned environments between stages, with practitioners guessing which configuration will best improve the current policy. This paper hands that job to the model, proposing an LLM-as-Environment-Engineer framework where the policy diagnoses its own weaknesses and proposes the next environment to train on.

70Reinforcement Learning
Back on Track

Back on Track

Diffusion large language models generate text in a way that does not fit cleanly into the reinforcement learning recipes built for autoregressive models, and training them to reason exposes two specific problems. Rewards are sparse, so a single terminal reward fails to guide intermediate generation steps, and policy updates sometimes drift toward unnatural trajectories rather than authentic generation paths. This paper proposes Process Aligned Policy Optimization to fix both.

71Reinforcement Learning
Beyond Scalar Rewards

Beyond Scalar Rewards

Reward models usually compress a judgment into a single scalar, but this paper argues human preferences are better captured as score distributions, and proposes Z-Reward, which internalizes reasoning into a predicted distribution before scoring. A large vision-language teacher does the reasoning-heavy judgment and is distilled into a compact student for efficient deployment, with the 27B teacher reaching 89.6% human-preference accuracy and the 9B student nearly matching it at 88.6%. Used as a reinforcement learning signal, it delivers a 41.3% net preference improvement over a supervised baseline, beating GRPO and other reward methods.

72Reinforcement Learning
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026