🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
200 papers · MultimodalClear filters →
Patch n' Pack: NaViT

Patch n' Pack: NaViT

A vision transformer handling any aspect ratio and resolution through sequence packing.

169Architecture
HyperDreamBooth

HyperDreamBooth

A smaller, faster, and more efficient version of DreamBooth for personalizing text-to-image models.

170Multimodal
AnimateDiff

AnimateDiff

Animates frozen text-to-image diffusion models via a plug-in motion modeling module.

171Multimodal
Generative Pretraining in Multimodality (Emu)

Generative Pretraining in Multimodality (Emu)

A transformer-based multimodal foundation model for generating images and text.

172Multimodal
Multimodal Generation with Frozen LLMs

Multimodal Generation with Frozen LLMs

Maps images to LLM token space enabling models like PaLM and GPT-4 to handle visual tasks without parameter updates.

173Multimodal
Physics-based Motion Retargeting in Real-Time

Physics-based Motion Retargeting in Real-Time

Uses RL to retarget motions from sparse human sensor data to characters of various morphologies.

174Multimodal
Computer Vision Through the Lens of Natural Language

Computer Vision Through the Lens of Natural Language

A modular approach solving CV problems by routing through LLM reasoning.

175Multimodal
DragDiffusion

DragDiffusion

Extends interactive point-based image editing to diffusion models.

176Multimodal
MotionGPT

MotionGPT

Generates consecutive human motions from multimodal control signals via LLM instructions.

177Multimodal
AudioPaLM

AudioPaLM

Fuses PaLM-2 and AudioLM into a multimodal architecture supporting speech understanding and generation.

178Multimodal
Voicebox

Voicebox

Meta's all-in-one generative speech model supporting 6 languages and many speech tasks in-context.

179Multimodal
TAPIR

TAPIR

Tracks any queried point on any physical surface throughout a video sequence faster than real-time.

180Multimodal
Tracking Everything Everywhere All at Once (OmniMotion)

Tracking Everything Everywhere All at Once (OmniMotion)

Test-time optimization for dense, long-range motion estimation.

181Multimodal
MusicGen

MusicGen

A simple and controllable model for music generation using a single-stage Transformer.

182Multimodal
Hierarchical Vision Transformer (Hiera)

Hierarchical Vision Transformer (Hiera)

Pretrains ViTs with MAE while removing unnecessary multi-stage complexity.

183Architecture
BiomedGPT

BiomedGPT

A unified biomedical GPT for vision, language, and multimodal tasks.

184Multimodal
MERT

MERT

An acoustic music understanding model with large-scale self-supervised training.

185Multimodal
Drag Your GAN (DragGAN)

Drag Your GAN (DragGAN)

Interactive point-based image manipulation on the generative image manifold.

186Multimodal
ImageBind

ImageBind

Meta's joint embedding across six modalities at once.

187Multimodal
InstructBLIP

InstructBLIP

Visual-language instruction tuning built on BLIP-2.

188Multimodal
MultiModal-GPT

MultiModal-GPT

A vision-language model for multi-round dialogue fine-tuned from OpenFlamingo.

189Multimodal
Shap-E

Shap-E

OpenAI's conditional generative model for 3D assets producing implicit functions.

190Multimodal
Track Anything

Track Anything

An interactive tool for video object tracking and segmentation built on Segment Anything.

191Multimodal
AudioGPT

AudioGPT

Connects ChatGPT with audio foundational models for speech, music, sound, and talking head tasks.

192Multimodal
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026