🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
224 papers · MultimodalClear filters →
Patch n' Pack: NaViT

Patch n' Pack: NaViT

A vision transformer handling any aspect ratio and resolution through sequence packing.

193Architecture
HyperDreamBooth

HyperDreamBooth

A smaller, faster, and more efficient version of DreamBooth for personalizing text-to-image models.

194Multimodal
AnimateDiff

AnimateDiff

Animates frozen text-to-image diffusion models via a plug-in motion modeling module.

195Multimodal
Generative Pretraining in Multimodality (Emu)

Generative Pretraining in Multimodality (Emu)

A transformer-based multimodal foundation model for generating images and text.

196Multimodal
Multimodal Generation with Frozen LLMs

Multimodal Generation with Frozen LLMs

Maps images to LLM token space enabling models like PaLM and GPT-4 to handle visual tasks without parameter updates.

197Multimodal
Physics-based Motion Retargeting in Real-Time

Physics-based Motion Retargeting in Real-Time

Uses RL to retarget motions from sparse human sensor data to characters of various morphologies.

198Multimodal
Computer Vision Through the Lens of Natural Language

Computer Vision Through the Lens of Natural Language

A modular approach solving CV problems by routing through LLM reasoning.

199Multimodal
DragDiffusion

DragDiffusion

Extends interactive point-based image editing to diffusion models.

200Multimodal
MotionGPT

MotionGPT

Generates consecutive human motions from multimodal control signals via LLM instructions.

201Multimodal
AudioPaLM

AudioPaLM

Fuses PaLM-2 and AudioLM into a multimodal architecture supporting speech understanding and generation.

202Multimodal
Voicebox

Voicebox

Meta's all-in-one generative speech model supporting 6 languages and many speech tasks in-context.

203Multimodal
Tracking Everything Everywhere All at Once (OmniMotion)

Tracking Everything Everywhere All at Once (OmniMotion)

Test-time optimization for dense, long-range motion estimation.

204Multimodal
MusicGen

MusicGen

A simple and controllable model for music generation using a single-stage Transformer.

205Multimodal
Hierarchical Vision Transformer (Hiera)

Hierarchical Vision Transformer (Hiera)

Pretrains ViTs with MAE while removing unnecessary multi-stage complexity.

206Architecture
BiomedGPT

BiomedGPT

A unified biomedical GPT for vision, language, and multimodal tasks.

207Multimodal
MERT

MERT

An acoustic music understanding model with large-scale self-supervised training.

208Multimodal
Drag Your GAN (DragGAN)

Drag Your GAN (DragGAN)

Interactive point-based image manipulation on the generative image manifold.

209Multimodal
ImageBind

ImageBind

Meta's joint embedding across six modalities at once.

210Multimodal
InstructBLIP

InstructBLIP

Visual-language instruction tuning built on BLIP-2.

211Multimodal
MultiModal-GPT

MultiModal-GPT

A vision-language model for multi-round dialogue fine-tuned from OpenFlamingo.

212Multimodal
Shap-E

Shap-E

OpenAI's conditional generative model for 3D assets producing implicit functions.

213Multimodal
Track Anything

Track Anything

An interactive tool for video object tracking and segmentation built on Segment Anything.

214Multimodal
AudioGPT

AudioGPT

Connects ChatGPT with audio foundational models for speech, music, sound, and talking head tasks.

215Multimodal
DataComp

DataComp

A multimodal dataset benchmark with 12.8B image-text pairs.

216Multimodal
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026