🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
200 papers · MultimodalClear filters →
Red Teaming Visual Language Models

Red Teaming Visual Language Models

Introduces the first dedicated red-teaming benchmark for VLMs, covering vulnerabilities unique to multimodal inputs.

121Evaluation
Lumiere

Lumiere

Google's Lumiere is a space-time diffusion model for text-to-video that generates the entire video duration in a single forward pass rather than cascading short clips.

122Multimodal
InseRF

InseRF

InseRF inserts brand-new 3D objects into Neural Radiance Field scenes from just a text prompt plus a 2D bounding box, without requiring any explicit 3D input.

123Multimodal
MagicVideo-V2

MagicVideo-V2

ByteDance's MagicVideo-V2 is an end-to-end text-to-video pipeline that stitches together four specialized modules into a high-fidelity generation system.

124Multimodal
Instruct-Imagen

Instruct-Imagen

Google's Instruct-Imagen is a multimodal instruction-tuned image generation model that generalizes across heterogeneous generation tasks, including unseen ones.

125Multimodal
From Gemini to Q-Star

From Gemini to Q-Star

A 300+-paper survey mapping the state of Generative AI and the research frontiers that followed the Gemini + rumored Q* news cycle.

126Multimodal
Gemini vs GPT-4V

Gemini vs GPT-4V

A qualitative side-by-side comparison of Gemini and GPT-4V across vision-language tasks, documenting systematic behavioral differences.

127Multimodal
VideoPoet

VideoPoet

Google Research's VideoPoet is a large language model for zero-shot video generation that treats video as just another token stream.

128Multimodal
AppAgent

AppAgent

Introduces an LLM-based multimodal agent that operates real smartphone apps through touch actions and screenshots.

129Multimodal
Audiobox

Audiobox

Meta's Audiobox is a unified flow-matching audio model that generates speech, sound effects, and music from natural-language and example prompts.

130Multimodal
Gemini 1.0

Gemini 1.0

Google launches Gemini 1.0, a multimodal family natively designed to reason across text, images, video, audio, and code from the ground up.

131Multimodal
Adversarial Diffusion Distillation (SDXL Turbo)

Adversarial Diffusion Distillation (SDXL Turbo)

Stability AI's ADD trains a student diffusion model that produces high-quality images in just 1-4 sampling steps.

132Training
Seamless

Seamless

Meta's Seamless is a family of models for end-to-end expressive, streaming cross-lingual speech communication.

133Multimodal
UniIR

UniIR

UniIR is a unified instruction-guided multimodal retriever that handles eight retrieval tasks across modalities with a single model.

134Multimodal
Translatotron 3

Translatotron 3

Google's Translatotron 3 performs speech-to-speech translation using only monolingual data - no parallel corpora required.

135Multimodal
Mirasol3B

Mirasol3B

Google's Mirasol3B is a multimodal model that decouples modalities into focused autoregressive components rather than forcing a single fused stream.

136Multimodal
Emu Video and Emu Edit

Emu Video and Emu Edit

Meta releases Emu Video and Emu Edit, a pair of diffusion models targeting controlled text-to-video generation and instruction-based image editing.

137Multimodal
JARVIS-1

JARVIS-1

An open-world multimodal agent for Minecraft that combines perception, planning, and memory into a self-improving system.

138Agents
MusicGen

MusicGen

Meta's MusicGen is a single-stage transformer LLM for music generation that operates over compressed discrete audio tokens.

139Multimodal
MetNet-3

MetNet-3

Google's MetNet-3 is a state-of-the-art neural weather model extending lead time and variable coverage well beyond prior observation-based models.

140Architecture
Matryoshka Diffusion Models

Matryoshka Diffusion Models

Apple introduces an end-to-end framework for high-resolution image and video synthesis that denoises across multiple resolutions jointly.

141Multimodal
Spectron

Spectron

Google's Spectron is a spoken-language model trained end-to-end on raw spectrograms rather than text or discrete audio tokens.

142Multimodal
CommonCanvas

CommonCanvas

Releases CommonCanvas, a text-to-image dataset composed entirely of Creative-Commons-licensed images.

143Data
Video Language Planning

Video Language Planning

Enables synthesizing complex long-horizon video plans for robotics via tree search over vision-language and text-to-video models.

144Multimodal
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026