🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
258 papers · SafetyClear filters →
Universal Adversarial LLM Attacks

Universal Adversarial LLM Attacks

Finds universal and transferable adversarial attacks that cause aligned models like ChatGPT and Bard to generate objectionable behaviors.

241Safety
Survey of Aligned LLMs

Survey of Aligned LLMs

A comprehensive overview of alignment approaches covering data, training, and evaluation.

242Safety
Llama 2

Llama 2

Meta's open-weight foundation model family with chat-tuned variants ranging from 7B to 70B parameters.

243Training
Challenges & Application of LLMs

Challenges & Application of LLMs

A comprehensive enumeration of open challenges and application domains for LLMs.

244Safety
FLASK

FLASK

Proposes fine-grained evaluation of LLMs decomposed into 12 alignment skill sets.

245Evaluation
Claude 2

Claude 2

Anthropic's second-generation LLM with a detailed model card on safety, alignment, and capabilities.

246Safety
Robots That Ask for Help

Robots That Ask for Help

A framework for calibrating LLM-based robot planners so they ask for help when uncertain.

247Robotics
An Overview of Catastrophic AI Risks

An Overview of Catastrophic AI Risks

Dan Hendrycks' comprehensive overview of catastrophic AI risk categories.

248Safety
Reliability of Watermarks for LLMs

Reliability of Watermarks for LLMs

Studies whether watermarks survive human rewriting and LLM paraphrasing.

249Safety
Concept Scrubbing in LLM (LEACE)

Concept Scrubbing in LLM (LEACE)

Least-squares Concept Erasure - erases a target concept from every layer of a neural network.

250Safety
LIMA

LIMA

Meta's 65B LLaMA fine-tuned on just 1,000 curated examples - showing alignment needs less data than believed.

251Training
The Larger They Are, the Harder They Fail

The Larger They Are, the Harder They Fail

Reveals inverse-scaling failures in LLM code generation.

252Code
LLM Explains Neurons in LLMs

LLM Explains Neurons in LLMs

OpenAI's automated interpretability pipeline using GPT-4 to explain GPT-2 neurons.

253Safety
Unfaithful Explanations in Chain-of-Thought Prompting

Unfaithful Explanations in Chain-of-Thought Prompting

Demonstrates CoT explanations can misrepresent the true reason for a model's prediction.

254Reasoning
Interpretable ML for Science with PySR

Interpretable ML for Science with PySR

An open-source library for practical symbolic regression in the sciences.

255Safety
Poisoning Language Models During Instruction Tuning

Poisoning Language Models During Instruction Tuning

Shows adversaries can poison LLMs via instruction tuning data.

256Training
MACHIAVELLI Benchmark

MACHIAVELLI Benchmark

A benchmark of 134 text-based Choose-Your-Own-Adventure games for measuring ethical trade-offs.

257Evaluation
Pythia

Pythia

EleutherAI's suite for analyzing LLMs across training and scaling.

258Training
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026