🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,762
Papers
176
Weekly issues
2023
Since
390 papers · 2023Clear filters →
The Hydra Effect

The Hydra Effect

DeepMind shows that language models exhibit self-repairing behavior when attention heads are ablated.

217Safety
Self-Check

Self-Check

Explores LLM capacity for self-checking on complex reasoning tasks requiring multi-step and non-linear thinking.

218Safety
Dynalang (Agents Model the World with Language)

Dynalang (Agents Model the World with Language)

UC Berkeley's Dynalang agent learns a multimodal world model predicting future text, video, and rewards.

219Agents
AutoRobotics-Zero

AutoRobotics-Zero

Discovers zero-shot adaptable robot policies from scratch, including the automatic discovery of Python control code.

220Robotics
Universal Adversarial LLM Attacks

Universal Adversarial LLM Attacks

Finds universal and transferable adversarial attacks that cause aligned models like ChatGPT and Bard to generate objectionable behaviors.

221Safety
RT-2

RT-2

Google DeepMind's end-to-end vision-language-action model that learns from both web and robotics data to control robots.

222Robotics
Med-PaLM Multimodal

Med-PaLM Multimodal

Introduces a generalist biomedical AI system and a new multimodal biomedical benchmark with 14 tasks.

223Multimodal
Tracking Anything in High Quality

Tracking Anything in High Quality

A framework for high-quality tracking-anything in videos combining segmentation and refinement.

224Multimodal
Foundation Models in Vision

Foundation Models in Vision

A comprehensive survey on foundational models for computer vision and their open research directions.

225Multimodal
L-Eval

L-Eval

A standardized evaluation suite for long-context language models.

226Evaluation
LoraHub

LoraHub

Enables efficient cross-task generalization via dynamic LoRA composition.

227Training
Survey of Aligned LLMs

Survey of Aligned LLMs

A comprehensive overview of alignment approaches covering data, training, and evaluation.

228Safety
WavJourney

WavJourney

Leverages LLMs to orchestrate audio generation models for compositional storytelling.

229Multimodal
FacTool

FacTool

A task- and domain-agnostic framework for factuality detection of LLM-generated text.

230Evaluation
Llama 2

Llama 2

Meta's open-weight foundation model family with chat-tuned variants ranging from 7B to 70B parameters.

231Training
How is ChatGPT's Behavior Changing Over Time?

How is ChatGPT's Behavior Changing Over Time?

Evaluates GPT-3.5 and GPT-4 over months to show significant behavioral drift in deployed systems.

232Evaluation
FlashAttention-2

FlashAttention-2

Tri Dao's follow-up to FlashAttention, dramatically improving attention throughput on modern GPUs.

233Efficiency
Measuring Faithfulness in Chain-of-Thought Reasoning

Measuring Faithfulness in Chain-of-Thought Reasoning

Anthropic's investigation into whether CoT reasoning actually reflects the model's internal decision process.

234Reasoning
Generative TV & Showrunner Agents

Generative TV & Showrunner Agents

Fable Studio's approach to generate episodic TV content using LLMs and multi-agent simulation.

235Agents
Challenges & Application of LLMs

Challenges & Application of LLMs

A comprehensive enumeration of open challenges and application domains for LLMs.

236Safety
Retentive Network (RetNet)

Retentive Network (RetNet)

Microsoft's proposed foundation architecture aiming to replace Transformer attention for LLMs.

237Architecture
Meta-Transformer

Meta-Transformer

A unified framework performing learning across 12 different modalities with a shared backbone.

238Architecture
Retrieve In-Context Examples for LLMs

Retrieve In-Context Examples for LLMs

A framework to iteratively train dense retrievers that identify high-quality in-context examples.

239Retrieval
FLASK

FLASK

Proposes fine-grained evaluation of LLMs decomposed into 12 alignment skill sets.

240Evaluation
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026