🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
277 papers · SafetyClear filters →
Back to Basics: Revisiting REINFORCE in RLHF

Back to Basics: Revisiting REINFORCE in RLHF

Cohere researchers argue that PPO is overkill for RLHF and that a simpler REINFORCE-style estimator works better in practice.

193Reinforcement Learning
TestGen-LLM

TestGen-LLM

Meta's TestGen-LLM uses LLMs to improve existing human-written tests - augmenting coverage rather than generating tests from scratch - while rigorously filtering LLM output for quality.

194Safety
Survey of LLMs

Survey of LLMs

A survey that maps the landscape of the three dominant LLM families - GPT, Llama, and PaLM - and the shared toolbox used to build and augment them.

195Evaluation
LLM Agents Can Autonomously Hack Websites

LLM Agents Can Autonomously Hack Websites

The paper shows GPT-4 agents with tool use and long context can autonomously exploit real websites, including performing blind SQL injection and schema extraction.

196Agents
LLM-based Multi-Agent Systems Survey

LLM-based Multi-Agent Systems Survey

A survey of the fast-growing LLM-based multi-agent systems space, covering both problem-solving applications and "world simulation" research.

197Agents
Advances in Multimodal LLMs

Advances in Multimodal LLMs

A comprehensive survey mapping design choices for architecture and training pipeline around multimodal large language models (MLLMs).

198Multimodal
Hallucination in LVLMs

Hallucination in LVLMs

A survey specifically scoped to hallucination in Large Vision-Language Models, a phenomenon that differs substantially from text-only LLM hallucination.

199Safety
WARM (Weighted Averaged Reward Models)

WARM (Weighted Averaged Reward Models)

WARM averages multiple fine-tuned reward models in weight space rather than ensembling their predictions, dramatically reducing RLHF inference cost.

200Reinforcement Learning
Red Teaming Visual Language Models

Red Teaming Visual Language Models

Introduces the first dedicated red-teaming benchmark for VLMs, covering vulnerabilities unique to multimodal inputs.

201Evaluation
Self-Rewarding Language Models

Self-Rewarding Language Models

Meta shows that an LLM can act as both actor and judge in its own alignment loop, generating training data without any external reward model.

202Reinforcement Learning
Sleeper Agents

Sleeper Agents

Anthropic shows that LLMs can be trained to act deceptively under specific triggers and that current safety training techniques fail to remove this hidden behavior.

203Safety
TrustLLM (Trustworthiness in LLMs)

TrustLLM (Trustworthiness in LLMs)

A 100+ page study that defines a principled framework for trustworthy LLMs and benchmarks 16 mainstream models across it.

204Evaluation
Persuasive Adversarial Prompts (PAP)

Persuasive Adversarial Prompts (PAP)

Turns 40 human-persuasion techniques into a taxonomy of jailbreaks that achieve 92% attack success on frontier models without any optimization.

205Safety
Mitigating Hallucination in LLMs

Mitigating Hallucination in LLMs

A survey cataloging 32 hallucination-mitigation techniques and organizing them into a practical taxonomy.

206Safety
From Gemini to Q-Star

From Gemini to Q-Star

A 300+-paper survey mapping the state of Generative AI and the research frontiers that followed the Gemini + rumored Q* news cycle.

207Multimodal
Exploiting Novel GPT-4 APIs

Exploiting Novel GPT-4 APIs

A red-team study of three newer GPT-4 API surfaces - fine-tuning, function calling, and knowledge retrieval - that reveals each introduces new attack vectors.

208Training
Fact Recalling in LLMs

Fact Recalling in LLMs

A mechanistic-interpretability study showing that early MLP layers function as a lookup table for factual recall.

209Safety
Survey of Reasoning with Foundation Models

Survey of Reasoning with Foundation Models

A comprehensive survey of reasoning with foundation models, covering tasks, methods, benchmarks, and future directions.

210Reasoning
Adversarial Attacks on GPT-4

Adversarial Attacks on GPT-4

Demonstrates that a trivially simple random-search procedure can jailbreak GPT-4 with high reliability.

211Safety
FunSearch

FunSearch

DeepMind's FunSearch uses LLMs as a mutation operator in an evolutionary loop to discover genuinely new mathematical knowledge.

212Safety
Weak-to-Strong Generalization

Weak-to-Strong Generalization

OpenAI's superalignment team shows that weak supervisors can still elicit capabilities from much stronger models - a first empirical signal for scalable oversight.

213Training
LLMs in Medicine

LLMs in Medicine

A comprehensive survey (300+ papers) of LLMs applied to medicine, from clinical tasks to biomedical research.

214Evaluation
Pearl

Pearl

Meta's Pearl is a production-ready reinforcement learning agent package designed for real-world deployment constraints.

215Agents
Llama Guard

Llama Guard

Meta's Llama Guard is a compact, instruction-tuned safety classifier built on Llama 2-7B for input/output moderation in conversational AI.

216Safety
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026