🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
224 papers · MultimodalClear filters →
CogAgent

CogAgent

Tsinghua's CogAgent is an 18B-parameter visual-language model purpose-built for GUI understanding and navigation, with unusually high input resolution.

145Evaluation
From Gemini to Q-Star

From Gemini to Q-Star

A 300+-paper survey mapping the state of Generative AI and the research frontiers that followed the Gemini + rumored Q* news cycle.

146Multimodal
Survey of Reasoning with Foundation Models

Survey of Reasoning with Foundation Models

A comprehensive survey of reasoning with foundation models, covering tasks, methods, benchmarks, and future directions.

147Reasoning
Gemini vs GPT-4V

Gemini vs GPT-4V

A qualitative side-by-side comparison of Gemini and GPT-4V across vision-language tasks, documenting systematic behavioral differences.

148Multimodal
VideoPoet

VideoPoet

Google Research's VideoPoet is a large language model for zero-shot video generation that treats video as just another token stream.

149Multimodal
AppAgent

AppAgent

Introduces an LLM-based multimodal agent that operates real smartphone apps through touch actions and screenshots.

150Multimodal
RAG for LLMs

RAG for LLMs

A broad survey of Retrieval-Augmented Generation research, organizing the rapidly growing literature into a coherent map.

151Retrieval
Audiobox

Audiobox

Meta's Audiobox is a unified flow-matching audio model that generates speech, sound effects, and music from natural-language and example prompts.

152Multimodal
Gemini 1.0

Gemini 1.0

Google launches Gemini 1.0, a multimodal family natively designed to reason across text, images, video, audio, and code from the ground up.

153Multimodal
Adversarial Diffusion Distillation (SDXL Turbo)

Adversarial Diffusion Distillation (SDXL Turbo)

Stability AI's ADD trains a student diffusion model that produces high-quality images in just 1-4 sampling steps.

154Training
Seamless

Seamless

Meta's Seamless is a family of models for end-to-end expressive, streaming cross-lingual speech communication.

155Safety
UniIR

UniIR

UniIR is a unified instruction-guided multimodal retriever that handles eight retrieval tasks across modalities with a single model.

156Multimodal
Mirasol3B

Mirasol3B

Google's Mirasol3B is a multimodal model that decouples modalities into focused autoregressive components rather than forcing a single fused stream.

157Multimodal
GAIA

GAIA

Meta's GAIA is a benchmark for general AI assistants that requires reasoning, multimodal handling, web browsing, and tool use to solve real-world questions.

158Agents
Emu Video and Emu Edit

Emu Video and Emu Edit

Meta releases Emu Video and Emu Edit, a pair of diffusion models targeting controlled text-to-video generation and instruction-based image editing.

159Multimodal
JARVIS-1

JARVIS-1

An open-world multimodal agent for Minecraft that combines perception, planning, and memory into a self-improving system.

160Agents
MusicGen

MusicGen

Meta's MusicGen is a single-stage transformer LLM for music generation that operates over compressed discrete audio tokens.

161Multimodal
MetNet-3

MetNet-3

Google's MetNet-3 is a state-of-the-art neural weather model extending lead time and variable coverage well beyond prior observation-based models.

162Architecture
Matryoshka Diffusion Models

Matryoshka Diffusion Models

Apple introduces an end-to-end framework for high-resolution image and video synthesis that denoises across multiple resolutions jointly.

163Multimodal
Spectron

Spectron

Google's Spectron is a spoken-language model trained end-to-end on raw spectrograms rather than text or discrete audio tokens.

164Multimodal
CommonCanvas

CommonCanvas

Releases CommonCanvas, a text-to-image dataset composed entirely of Creative-Commons-licensed images.

165Data
Video Language Planning

Video Language Planning

Enables synthesizing complex long-horizon video plans for robotics via tree search over vision-language and text-to-video models.

166Multimodal
UniSim (Universal Simulator)

UniSim (Universal Simulator)

Google's UniSim learns a universal generative simulator of real-world interactions from diverse video + action data.

167Robotics
The Dawn of LMMs (GPT-4V Deep Dive)

The Dawn of LMMs (GPT-4V Deep Dive)

Microsoft's exhaustive 166-page analysis of GPT-4V's capabilities and limitations.

168Multimodal
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026