🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
390 papers · 2023Clear filters →
ConvNets Match Vision Transformers

ConvNets Match Vision Transformers

DeepMind shows that strong ConvNet architectures pretrained at scale match ViTs on ImageNet performance at comparable compute.

97Training
CommonCanvas

CommonCanvas

Releases CommonCanvas, a text-to-image dataset composed entirely of Creative-Commons-licensed images.

98Data
Managing AI Risks (Bengio, Hinton, et al.)

Managing AI Risks (Bengio, Hinton, et al.)

A high-profile position paper by leading AI researchers laying out risks from upcoming advanced AI systems.

99Safety
Branch-Solve-Merge (BSM)

Branch-Solve-Merge (BSM)

BSM decomposes LLM tasks into parallel sub-tasks via three LLM-programmed modules: branch, solve, and merge.

100Agents
Llemma

Llemma

Llemma is an open LLM for mathematics built via continued pretraining of Code Llama on the Proof-Pile-2 dataset.

101Data
LLMs for Software Engineering

LLMs for Software Engineering

A comprehensive survey of LLMs for software engineering covering models, tasks, evaluation, and open challenges.

102Code
Self-RAG

Self-RAG

Self-RAG trains an LM to adaptively retrieve, generate, and self-critique using special reflection tokens.

103Retrieval
RAG for Long-Form QA

RAG for Long-Form QA

Explores retrieval-augmented LMs specifically on long-form question answering, where RAG failures are more subtle.

104Retrieval
GenBench

GenBench

A Nature Machine Intelligence paper framework for characterizing and understanding generalization research in NLP.

105Evaluation
LLM Self-Explanations

LLM Self-Explanations

Investigates whether LLMs can generate useful feature-attribution explanations for their own outputs.

106Safety
OpenAgents

OpenAgents

An open platform for running and hosting real-world language agents, including three distinct agent types.

107Agents
Eliciting Human Preferences with LLMs

Eliciting Human Preferences with LLMs

Anthropic uses LLMs to guide the task-specification process, eliciting user intent through natural-language dialogue.

108Reinforcement Learning
AutoMix

AutoMix

AutoMix routes queries between LLMs of different sizes based on smaller-model confidence, saving cost without sacrificing quality.

109Efficiency
Video Language Planning

Video Language Planning

Enables synthesizing complex long-horizon video plans for robotics via tree search over vision-language and text-to-video models.

110Multimodal
Ring Attention

Ring Attention

UC Berkeley's Ring Attention scales transformer context to 100M+ tokens by distributing blockwise self-attention across devices in a ring topology.

111Memory
UniSim (Universal Simulator)

UniSim (Universal Simulator)

Google's UniSim learns a universal generative simulator of real-world interactions from diverse video + action data.

112Robotics
Survey on Factuality in LLMs

Survey on Factuality in LLMs

A survey covering evaluation and enhancement techniques for LLM factuality.

113Evaluation
Hypothesis Search (LLMs Can Learn Rules)

Hypothesis Search (LLMs Can Learn Rules)

A two-stage framework where the LLM learns a rule library for reasoning.

114Reasoning
Meta Chain-of-Thought Prompting (Meta-CoT)

Meta Chain-of-Thought Prompting (Meta-CoT)

A generalizable CoT framework that selects domain-appropriate reasoning patterns for the task at hand.

115Reasoning
LLMs for Healthcare Survey

LLMs for Healthcare Survey

A comprehensive overview of LLMs applied to the healthcare domain.

116Evaluation
RECOMP (Retrieval-Augmented LMs with Compressors)

RECOMP (Retrieval-Augmented LMs with Compressors)

Proposes two compression approaches to shrink retrieved documents before in-context use.

117Retrieval
InstructRetro

InstructRetro

NVIDIA introduces Retro 48B, the largest LLM pretrained with retrieval at the time.

118Training
MemWalker

MemWalker

MemWalker treats the LLM as an interactive agent that traverses a tree-structured summary of long text.

119Memory
FireAct (Language Agent Fine-tuning)

FireAct (Language Agent Fine-tuning)

Explores fine-tuning LLMs specifically for language-agent use, demonstrating consistent gains over prompting alone.

120Agents
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026