🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
242 papers · Reinforcement LearningClear filters →
RLHF Workflow

RLHF Workflow

provides an easily reproducible recipe for online iterative RLHF; discusses theoretical insights and algorithmic principles of online iterative RLHF and practical implementation.

217Reinforcement Learning
Self-Play Preference Optimization

Self-Play Preference Optimization

proposes a self-play-based method for aligning language models; this optimation procedure treats the problem as a constant-sum two-player game to identify the Nash equilibrium policy; it addresses the shortcomings of DPO and IPO and effectively increases the log-likelihood of chose responses and decreases the rejected ones; SPPO outperforms DPO and IPO on MT-Bench and the Open LLM Leaderboard.

218Reinforcement Learning
Gemma

Gemma

Google DeepMind releases Gemma, a family of open models (2B and 7B) built from the same research stack as Gemini and shipped with both base and instruction-tuned variants.

219Reinforcement Learning
Back to Basics: Revisiting REINFORCE in RLHF

Back to Basics: Revisiting REINFORCE in RLHF

Cohere researchers argue that PPO is overkill for RLHF and that a simpler REINFORCE-style estimator works better in practice.

220Reinforcement Learning
WARM (Weighted Averaged Reward Models)

WARM (Weighted Averaged Reward Models)

WARM averages multiple fine-tuned reward models in weight space rather than ensembling their predictions, dramatically reducing RLHF inference cost.

221Reinforcement Learning
Self-Rewarding Language Models

Self-Rewarding Language Models

Meta shows that an LLM can act as both actor and judge in its own alignment loop, generating training data without any external reward model.

222Reinforcement Learning
ReFT (Reinforced Fine-Tuning)

ReFT (Reinforced Fine-Tuning)

ByteDance's ReFT enhances LLM reasoning by combining supervised fine-tuning with online RL that samples alternative reasoning paths, without a learned reward model.

223Reasoning
Self-Play Fine-Tuning (SPIN)

Self-Play Fine-Tuning (SPIN)

SPIN shows that a supervised fine-tuned LLM can keep improving via self-play alone, without any additional human annotations.

224Training
Pearl

Pearl

Meta's Pearl is a production-ready reinforcement learning agent package designed for real-world deployment constraints.

225Agents
KTO (Kahneman-Tversky Optimization)

KTO (Kahneman-Tversky Optimization)

Contextual AI introduces KTO, an alignment objective derived from prospect theory that works with binary "good/bad" signals instead of preference pairs.

226Reinforcement Learning
TÜLU 2

TÜLU 2

Allen AI's TÜLU 2 is a suite of improved open instruction-tuned LLMs and an accompanying study of adaptation best practices.

227Training
Zephyr

Zephyr

Hugging Face's Zephyr-7B is a 7B parameter LLM whose chat performance rivals much larger chat models aligned with human feedback.

228Reinforcement Learning
Eliciting Human Preferences with LLMs

Eliciting Human Preferences with LLMs

Anthropic uses LLMs to guide the task-specification process, eliciting user intent through natural-language dialogue.

229Reinforcement Learning
LLaVA-RLHF

LLaVA-RLHF

Adapts factually augmented RLHF to aligning large multimodal models, reducing hallucination without falling into reward-hacking pitfalls.

230Reinforcement Learning
RLAIF (Scaling RLHF with AI Feedback)

RLAIF (Scaling RLHF with AI Feedback)

Google compares RLHF with RLAIF (Reinforcement Learning from AI Feedback) to test whether AI preferences can replace human preferences.

231Reinforcement Learning
Open Problems and Limitations of RLHF

Open Problems and Limitations of RLHF

A comprehensive survey of open problems and fundamental limitations of RLHF as an alignment approach.

232Reinforcement Learning
Survey of Aligned LLMs

Survey of Aligned LLMs

A comprehensive overview of alignment approaches covering data, training, and evaluation.

233Safety
Llama 2

Llama 2

Meta's open-weight foundation model family with chat-tuned variants ranging from 7B to 70B parameters.

234Training
Secrets of RLHF in LLMs

Secrets of RLHF in LLMs

A deep investigation into RLHF with a focus on the inner workings of PPO, including open-source code.

235Reinforcement Learning
Elastic Decision Transformer

Elastic Decision Transformer

An advance over Decision Transformers that enables trajectory stitching at inference time.

236Reinforcement Learning
InterCode

InterCode

A framework treating interactive coding as a reinforcement learning environment.

237Reinforcement Learning
SequenceMatch

SequenceMatch

Formulates sequence generation as imitation learning, enabling backtracking via a backspace action.

238Reinforcement Learning
AlphaDev

AlphaDev

DeepMind's deep RL agent discovering faster sorting algorithms from scratch, now in LLVM.

239Reinforcement Learning
Fine-Grained RLHF

Fine-Grained RLHF

Trains LMs with segment-level human feedback rather than whole-response preferences.

240Reinforcement Learning
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026