🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
173 papers · Reinforcement LearningClear filters →
Discovering Preference Optimization Algorithms with LLMs

Discovering Preference Optimization Algorithms with LLMs

proposes LLM-driven objective discovery of state-of-the-art preference optimization; no human intervention is used and an LLM is prompted to propose and implement the preference optimization loss functions based on previously evaluated performance metrics; discovers an algorithm that adaptively combined logistic and exponential losses.

145Reinforcement Learning
SaySelf

SaySelf

a training framework to teach LLMs to express more accurate fine-grained confidence estimates and self-reflective rationales; it performs supervised finetuning on a dataset that contains summaries of the difference between multiple reasoning chains; reinforcement learning is then applied to calibrate confidence estimates, encouraging the LLM to produce accurate, high-confidence predictions and penalize overconfidence in erroneous outputs.

146Reinforcement Learning
SimPO

SimPO

a simpler and more effective approach for preference optimization with a reference-free reward; uses the average log probability of a sequence as an implicit reward (i.e., no reference model required) which makes it more compute and memory efficient; demonstrates that it outperforms existing approaches like DPO and claims to produce the strongest 8B open-source model.

147Reinforcement Learning
RLHF Workflow

RLHF Workflow

provides an easily reproducible recipe for online iterative RLHF; discusses theoretical insights and algorithmic principles of online iterative RLHF and practical implementation.

148Reinforcement Learning
Self-Play Preference Optimization

Self-Play Preference Optimization

proposes a self-play-based method for aligning language models; this optimation procedure treats the problem as a constant-sum two-player game to identify the Nash equilibrium policy; it addresses the shortcomings of DPO and IPO and effectively increases the log-likelihood of chose responses and decreases the rejected ones; SPPO outperforms DPO and IPO on MT-Bench and the Open LLM Leaderboard.

149Reinforcement Learning
Gemma

Gemma

Google DeepMind releases Gemma, a family of open models (2B and 7B) built from the same research stack as Gemini and shipped with both base and instruction-tuned variants.

150Reinforcement Learning
Back to Basics: Revisiting REINFORCE in RLHF

Back to Basics: Revisiting REINFORCE in RLHF

Cohere researchers argue that PPO is overkill for RLHF and that a simpler REINFORCE-style estimator works better in practice.

151Reinforcement Learning
WARM (Weighted Averaged Reward Models)

WARM (Weighted Averaged Reward Models)

WARM averages multiple fine-tuned reward models in weight space rather than ensembling their predictions, dramatically reducing RLHF inference cost.

152Reinforcement Learning
Self-Rewarding Language Models

Self-Rewarding Language Models

Meta shows that an LLM can act as both actor and judge in its own alignment loop, generating training data without any external reward model.

153Reinforcement Learning
ReFT (Reinforced Fine-Tuning)

ReFT (Reinforced Fine-Tuning)

ByteDance's ReFT enhances LLM reasoning by combining supervised fine-tuning with online RL that samples alternative reasoning paths, without a learned reward model.

154Reasoning
Self-Play Fine-Tuning (SPIN)

Self-Play Fine-Tuning (SPIN)

SPIN shows that a supervised fine-tuned LLM can keep improving via self-play alone, without any additional human annotations.

155Training
Pearl

Pearl

Meta's Pearl is a production-ready reinforcement learning agent package designed for real-world deployment constraints.

156Agents
KTO (Kahneman-Tversky Optimization)

KTO (Kahneman-Tversky Optimization)

Contextual AI introduces KTO, an alignment objective derived from prospect theory that works with binary "good/bad" signals instead of preference pairs.

157Reinforcement Learning
TÜLU 2

TÜLU 2

Allen AI's TÜLU 2 is a suite of improved open instruction-tuned LLMs and an accompanying study of adaptation best practices.

158Training
Zephyr

Zephyr

Hugging Face's Zephyr-7B is a 7B parameter LLM whose chat performance rivals much larger chat models aligned with human feedback.

159Reinforcement Learning
Eliciting Human Preferences with LLMs

Eliciting Human Preferences with LLMs

Anthropic uses LLMs to guide the task-specification process, eliciting user intent through natural-language dialogue.

160Reinforcement Learning
LLaVA-RLHF

LLaVA-RLHF

Adapts factually augmented RLHF to aligning large multimodal models, reducing hallucination without falling into reward-hacking pitfalls.

161Reinforcement Learning
RLAIF (Scaling RLHF with AI Feedback)

RLAIF (Scaling RLHF with AI Feedback)

Google compares RLHF with RLAIF (Reinforcement Learning from AI Feedback) to test whether AI preferences can replace human preferences.

162Reinforcement Learning
Open Problems and Limitations of RLHF

Open Problems and Limitations of RLHF

A comprehensive survey of open problems and fundamental limitations of RLHF as an alignment approach.

163Reinforcement Learning
Survey of Aligned LLMs

Survey of Aligned LLMs

A comprehensive overview of alignment approaches covering data, training, and evaluation.

164Safety
Llama 2

Llama 2

Meta's open-weight foundation model family with chat-tuned variants ranging from 7B to 70B parameters.

165Training
Secrets of RLHF in LLMs

Secrets of RLHF in LLMs

A deep investigation into RLHF with a focus on the inner workings of PPO, including open-source code.

166Reinforcement Learning
Elastic Decision Transformer

Elastic Decision Transformer

An advance over Decision Transformers that enables trajectory stitching at inference time.

167Reinforcement Learning
InterCode

InterCode

A framework treating interactive coding as a reinforcement learning environment.

168Reinforcement Learning
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026