🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
277 papers · SafetyClear filters →
Humpback (Self-Alignment with Instruction Backtranslation)

Humpback (Self-Alignment with Instruction Backtranslation)

Meta's Humpback automatically generates instruction-tuning data by back-translating web text into plausible instructions.

241Safety
Shepherd

Shepherd

Meta's Shepherd is a 7B language model specifically tuned to critique model outputs and suggest refinements.

242Safety
Political Biases in NLP Models

Political Biases in NLP Models

Develops methods to measure political and media biases in LLMs and their downstream effects.

243Safety
Studying LLM Generalization with Influence Functions

Studying LLM Generalization with Influence Functions

Anthropic scales influence functions to LLMs up to 52B parameters to investigate generalization patterns.

244Safety
Synthetic Data Reduces Sycophancy

Synthetic Data Reduces Sycophancy

Google shows that fine-tuning on simple synthetic data can significantly reduce LLM sycophancy.

245Data
Trustworthy LLMs

Trustworthy LLMs

Presents a comprehensive framework of categories for assessing LLM trustworthiness.

246Safety
Open Problems and Limitations of RLHF

Open Problems and Limitations of RLHF

A comprehensive survey of open problems and fundamental limitations of RLHF as an alignment approach.

247Reinforcement Learning
The Hydra Effect

The Hydra Effect

DeepMind shows that language models exhibit self-repairing behavior when attention heads are ablated.

248Safety
Self-Check

Self-Check

Explores LLM capacity for self-checking on complex reasoning tasks requiring multi-step and non-linear thinking.

249Safety
Universal Adversarial LLM Attacks

Universal Adversarial LLM Attacks

Finds universal and transferable adversarial attacks that cause aligned models like ChatGPT and Bard to generate objectionable behaviors.

250Safety
Survey of Aligned LLMs

Survey of Aligned LLMs

A comprehensive overview of alignment approaches covering data, training, and evaluation.

251Safety
Llama 2

Llama 2

Meta's open-weight foundation model family with chat-tuned variants ranging from 7B to 70B parameters.

252Training
How is ChatGPT's Behavior Changing Over Time?

How is ChatGPT's Behavior Changing Over Time?

Evaluates GPT-3.5 and GPT-4 over months to show significant behavioral drift in deployed systems.

253Evaluation
Measuring Faithfulness in Chain-of-Thought Reasoning

Measuring Faithfulness in Chain-of-Thought Reasoning

Anthropic's investigation into whether CoT reasoning actually reflects the model's internal decision process.

254Reasoning
Challenges & Application of LLMs

Challenges & Application of LLMs

A comprehensive enumeration of open challenges and application domains for LLMs.

255Safety
FLASK

FLASK

Proposes fine-grained evaluation of LLMs decomposed into 12 alignment skill sets.

256Evaluation
Claude 2

Claude 2

Anthropic's second-generation LLM with a detailed model card on safety, alignment, and capabilities.

257Safety
Robots That Ask for Help

Robots That Ask for Help

A framework for calibrating LLM-based robot planners so they ask for help when uncertain.

258Robotics
An Overview of Catastrophic AI Risks

An Overview of Catastrophic AI Risks

Dan Hendrycks' comprehensive overview of catastrophic AI risk categories.

259Safety
LMFlow

LMFlow

An extensible and lightweight toolkit for fine-tuning and inference of large foundation models.

260Training
Reliability of Watermarks for LLMs

Reliability of Watermarks for LLMs

Studies whether watermarks survive human rewriting and LLM paraphrasing.

261Safety
Concept Scrubbing in LLM (LEACE)

Concept Scrubbing in LLM (LEACE)

Least-squares Concept Erasure - erases a target concept from every layer of a neural network.

262Safety
Direct Preference Optimization (DPO)

Direct Preference Optimization (DPO)

Rafailov et al.'s simpler alternative to RLHF that rivals full RL-based alignment.

263Reinforcement Learning
LIMA

LIMA

Meta's 65B LLaMA fine-tuned on just 1,000 curated examples - showing alignment needs less data than believed.

264Training
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026