🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
390 papers · 2023Clear filters →
Struc-Bench (LLMs for Structured Data)

Struc-Bench (LLMs for Structured Data)

Studies how LLMs handle complex structured-data generation and proposes a structure-aware fine-tuning method.

145Evaluation
LMSYS-Chat-1M

LMSYS-Chat-1M

LMSYS releases a large-scale dataset of 1 million real-world LLM conversations collected from the Vicuna demo and Chatbot Arena.

146Data
Language Modeling Is Compression

Language Modeling Is Compression

DeepMind empirically revisits the theoretical equivalence between prediction and compression, applied to modern LLMs.

147Efficiency
Compositional Foundation Models (HiP)

Compositional Foundation Models (HiP)

Proposes foundation models that compose multiple expert foundation models trained on different modalities to solve long-horizon goals.

148Architecture
OWL (LLMs for IT Operations)

OWL (LLMs for IT Operations)

Proposes OWL, an LLM specialized for IT operations through self-instruct fine-tuning on IT-specific tasks.

149Evaluation
KOSMOS-2.5

KOSMOS-2.5

Microsoft's KOSMOS-2.5 is a multimodal model purpose-built for "machine reading" of text-intensive images.

150Multimodal
Textbooks Are All You Need II (phi-1.5)

Textbooks Are All You Need II (phi-1.5)

Microsoft's phi-1.5 demonstrates that a 1.3B model trained on "textbook-quality" synthetic data rivals much larger models on reasoning.

151Data
The Rise and Potential of LLM-Based Agents

The Rise and Potential of LLM-Based Agents

A comprehensive survey of LLM-based agents covering construction, capability, and societal implications.

152Agents
EvoDiff

EvoDiff

Microsoft's EvoDiff combines evolutionary-scale protein data with diffusion models for controllable protein generation in sequence space.

153Architecture
Rewindable Auto-regressive INference (RAIN)

Rewindable Auto-regressive INference (RAIN)

Shows that unaligned LLMs can produce aligned responses at inference time via self-evaluation and rewinding.

154Safety
Robot Parkour Learning

Robot Parkour Learning

Stanford's Robot Parkour system learns end-to-end vision-based parkour policies that transfer to a quadrupedal robot.

155Robotics
Hallucination Survey (Early)

Hallucination Survey (Early)

Classifies hallucination phenomena in LLMs and catalogs evaluation criteria and mitigation strategies.

156Safety
Agents Library

Agents Library

An open-source library for building autonomous language agents with first-class support for planning, memory, tools, and multi-agent communication.

157Agents
Radiology-Llama 2

Radiology-Llama 2

A Llama 2-based LLM specialized for radiology report generation.

158Training
ChatDev (Communicative Agents for Software Development)

ChatDev (Communicative Agents for Software Development)

ChatDev is a virtual chat-powered software company where LLM agents take on roles in a waterfall-model dev process.

159Agents
MAmmoTH

MAmmoTH

An open-source LLM family specialized for general mathematical problem solving.

160Reasoning
Transformers as Support Vector Machines

Transformers as Support Vector Machines

A theoretical paper establishing a formal connection between self-attention optimization and hard-margin SVM problems.

161Training
RLAIF (Scaling RLHF with AI Feedback)

RLAIF (Scaling RLHF with AI Feedback)

Google compares RLHF with RLAIF (Reinforcement Learning from AI Feedback) to test whether AI preferences can replace human preferences.

162Reinforcement Learning
GPT Solves Math Problems Without a Calculator

GPT Solves Math Problems Without a Calculator

Demonstrates that with sufficient training data, even a small language model can perform accurate multi-digit arithmetic.

163Reasoning
OPRO (LLMs as Optimizers)

OPRO (LLMs as Optimizers)

DeepMind's OPRO uses LLMs as general-purpose optimizers over natural-language-described problems.

164Agents
ImageBind-LLM

ImageBind-LLM

Shanghai AI Lab's ImageBind-LLM brings six-modality understanding to LLMs via the ImageBind joint embedding space.

165Multimodal
Explaining Grokking

Explaining Grokking

DeepMind advances our understanding of grokking, predicting and confirming two novel phenomena that test their theory.

166Safety
Overview of AI Deception

Overview of AI Deception

A survey cataloguing empirical examples of AI systems exhibiting deceptive behavior.

167Safety
FLM-101B

FLM-101B

A 101B parameter open LLM trainable on a $100K budget through a growth-based training strategy.

168Training
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026