🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
277 papers · SafetyClear filters →
KTO (Kahneman-Tversky Optimization)

KTO (Kahneman-Tversky Optimization)

Contextual AI introduces KTO, an alignment objective derived from prospect theory that works with binary "good/bad" signals instead of preference pairs.

217Reinforcement Learning
Seamless

Seamless

Meta's Seamless is a family of models for end-to-end expressive, streaming cross-lingual speech communication.

218Safety
Safe Deployment of Generative AI (Nature)

Safe Deployment of Generative AI (Nature)

A Nature correspondence arguing that medical professionals - not commercial interests - must drive the development and deployment of generative AI in medicine.

219Safety
Translatotron 3

Translatotron 3

Google's Translatotron 3 performs speech-to-speech translation using only monolingual data - no parallel corpora required.

220Data
Fine-Tuning LLMs for Factuality

Fine-Tuning LLMs for Factuality

Stanford fine-tunes LLMs for factuality without any human labels by using automatically generated preference signals.

221Training
MART (Multi-round Automatic Red-Teaming)

MART (Multi-round Automatic Red-Teaming)

Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.

222Safety
LLMs Can Deceive Users (Trading Agent)

LLMs Can Deceive Users (Trading Agent)

Apollo Research shows that a helpful, honest LLM stock-trading agent can spontaneously deceive users under pressure.

223Agents
Hallucination in LLMs Survey

Hallucination in LLMs Survey

A comprehensive survey of hallucination in LLMs, covering taxonomy, causes, evaluation, and mitigation.

224Safety
Evaluating LLMs Survey

Evaluating LLMs Survey

A comprehensive survey of LLM evaluation covering benchmarks, methodologies, and open problems.

225Evaluation
Zephyr

Zephyr

Hugging Face's Zephyr-7B is a 7B parameter LLM whose chat performance rivals much larger chat models aligned with human feedback.

226Reinforcement Learning
Managing AI Risks (Bengio, Hinton, et al.)

Managing AI Risks (Bengio, Hinton, et al.)

A high-profile position paper by leading AI researchers laying out risks from upcoming advanced AI systems.

227Safety
LLM Self-Explanations

LLM Self-Explanations

Investigates whether LLMs can generate useful feature-attribution explanations for their own outputs.

228Safety
LLMs for Healthcare Survey

LLMs for Healthcare Survey

A comprehensive overview of LLMs applied to the healthcare domain.

229Evaluation
LLaVA-RLHF

LLaVA-RLHF

Adapts factually augmented RLHF to aligning large multimodal models, reducing hallucination without falling into reward-hacking pitfalls.

230Reinforcement Learning
LLM Alignment Survey

LLM Alignment Survey

A comprehensive survey of LLM alignment research spanning theoretical foundations to adversarial pressure.

231Safety
MentaLLaMA

MentaLLaMA

An open-source LLM family specialized for interpretable mental-health analysis on social media.

232Safety
Chain-of-Verification (CoVe)

Chain-of-Verification (CoVe)

Meta's Chain-of-Verification adds a "deliberation" step where the LLM fact-checks its own draft before finalizing.

233Safety
The Rise and Potential of LLM-Based Agents

The Rise and Potential of LLM-Based Agents

A comprehensive survey of LLM-based agents covering construction, capability, and societal implications.

234Agents
Rewindable Auto-regressive INference (RAIN)

Rewindable Auto-regressive INference (RAIN)

Shows that unaligned LLMs can produce aligned responses at inference time via self-evaluation and rewinding.

235Safety
Hallucination Survey (Early)

Hallucination Survey (Early)

Classifies hallucination phenomena in LLMs and catalogs evaluation criteria and mitigation strategies.

236Safety
RLAIF (Scaling RLHF with AI Feedback)

RLAIF (Scaling RLHF with AI Feedback)

Google compares RLHF with RLAIF (Reinforcement Learning from AI Feedback) to test whether AI preferences can replace human preferences.

237Reinforcement Learning
Explaining Grokking

Explaining Grokking

DeepMind advances our understanding of grokking, predicting and confirming two novel phenomena that test their theory.

238Safety
Overview of AI Deception

Overview of AI Deception

A survey cataloguing empirical examples of AI systems exhibiting deceptive behavior.

239Safety
LLMs for Illicit Purposes

LLMs for Illicit Purposes

A survey cataloguing threats and vulnerabilities arising from LLM deployment.

240Safety
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026