🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
258 papers · SafetyClear filters →
LLMs in Medicine

LLMs in Medicine

A comprehensive survey (300+ papers) of LLMs applied to medicine, from clinical tasks to biomedical research.

217Evaluation
Llama Guard

Llama Guard

Meta's Llama Guard is a compact, instruction-tuned safety classifier built on Llama 2-7B for input/output moderation in conversational AI.

218Safety
KTO (Kahneman-Tversky Optimization)

KTO (Kahneman-Tversky Optimization)

Contextual AI introduces KTO, an alignment objective derived from prospect theory that works with binary "good/bad" signals instead of preference pairs.

219Reinforcement Learning
Safe Deployment of Generative AI (Nature)

Safe Deployment of Generative AI (Nature)

A Nature correspondence arguing that medical professionals - not commercial interests - must drive the development and deployment of generative AI in medicine.

220Safety
Fine-Tuning LLMs for Factuality

Fine-Tuning LLMs for Factuality

Stanford fine-tunes LLMs for factuality without any human labels by using automatically generated preference signals.

221Training
MART (Multi-round Automatic Red-Teaming)

MART (Multi-round Automatic Red-Teaming)

Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.

222Safety
LLMs Can Deceive Users (Trading Agent)

LLMs Can Deceive Users (Trading Agent)

Apollo Research shows that a helpful, honest LLM stock-trading agent can spontaneously deceive users under pressure.

223Agents
Hallucination in LLMs Survey

Hallucination in LLMs Survey

A comprehensive survey of hallucination in LLMs, covering taxonomy, causes, evaluation, and mitigation.

224Safety
Zephyr

Zephyr

Hugging Face's Zephyr-7B is a 7B parameter LLM whose chat performance rivals much larger chat models aligned with human feedback.

225Reinforcement Learning
Managing AI Risks (Bengio, Hinton, et al.)

Managing AI Risks (Bengio, Hinton, et al.)

A high-profile position paper by leading AI researchers laying out risks from upcoming advanced AI systems.

226Safety
LLM Self-Explanations

LLM Self-Explanations

Investigates whether LLMs can generate useful feature-attribution explanations for their own outputs.

227Safety
LLMs Represent Space and Time

LLMs Represent Space and Time

MIT researchers find that LLMs internally encode linear representations of space and time across multiple scales.

228Safety
LLaVA-RLHF

LLaVA-RLHF

Adapts factually augmented RLHF to aligning large multimodal models, reducing hallucination without falling into reward-hacking pitfalls.

229Reinforcement Learning
LLM Alignment Survey

LLM Alignment Survey

A comprehensive survey of LLM alignment research spanning theoretical foundations to adversarial pressure.

230Safety
MentaLLaMA

MentaLLaMA

An open-source LLM family specialized for interpretable mental-health analysis on social media.

231Safety
Rewindable Auto-regressive INference (RAIN)

Rewindable Auto-regressive INference (RAIN)

Shows that unaligned LLMs can produce aligned responses at inference time via self-evaluation and rewinding.

232Safety
Hallucination Survey (Early)

Hallucination Survey (Early)

Classifies hallucination phenomena in LLMs and catalogs evaluation criteria and mitigation strategies.

233Safety
Explaining Grokking

Explaining Grokking

DeepMind advances our understanding of grokking, predicting and confirming two novel phenomena that test their theory.

234Safety
Overview of AI Deception

Overview of AI Deception

A survey cataloguing empirical examples of AI systems exhibiting deceptive behavior.

235Safety
LLMs for Illicit Purposes

LLMs for Illicit Purposes

A survey cataloguing threats and vulnerabilities arising from LLM deployment.

236Safety
Studying LLM Generalization with Influence Functions

Studying LLM Generalization with Influence Functions

Anthropic scales influence functions to LLMs up to 52B parameters to investigate generalization patterns.

237Safety
Synthetic Data Reduces Sycophancy

Synthetic Data Reduces Sycophancy

Google shows that fine-tuning on simple synthetic data can significantly reduce LLM sycophancy.

238Data
Trustworthy LLMs

Trustworthy LLMs

Presents a comprehensive framework of categories for assessing LLM trustworthiness.

239Safety
The Hydra Effect

The Hydra Effect

DeepMind shows that language models exhibit self-repairing behavior when attention heads are ablated.

240Safety
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026