🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
200 papers · MultimodalClear filters →
UniSim (Universal Simulator)

UniSim (Universal Simulator)

Google's UniSim learns a universal generative simulator of real-world interactions from diverse video + action data.

145Robotics
The Dawn of LMMs (GPT-4V Deep Dive)

The Dawn of LMMs (GPT-4V Deep Dive)

Microsoft's exhaustive 166-page analysis of GPT-4V's capabilities and limitations.

146Multimodal
KOSMOS-G

KOSMOS-G

Microsoft's KOSMOS-G extends zero-shot image generation to multi-image vision-language input.

147Multimodal
LLaVA-RLHF

LLaVA-RLHF

Adapts factually augmented RLHF to aligning large multimodal models, reducing hallucination without falling into reward-hacking pitfalls.

148Reinforcement Learning
KOSMOS-2.5

KOSMOS-2.5

Microsoft's KOSMOS-2.5 is a multimodal model purpose-built for "machine reading" of text-intensive images.

149Multimodal
ImageBind-LLM

ImageBind-LLM

Shanghai AI Lab's ImageBind-LLM brings six-modality understanding to LLMs via the ImageBind joint embedding space.

150Multimodal
LLaSM (Large Language and Speech Model)

LLaSM (Large Language and Speech Model)

A combined language-and-speech model trained with cross-modal conversational abilities.

151Multimodal
MVDream

MVDream

ByteDance's MVDream is a multi-view diffusion model that generates geometrically consistent images from multiple viewpoints given a text prompt.

152Multimodal
Nougat

Nougat

Meta's Nougat is a visual transformer for "Neural Optical Understanding for Academic documents" that converts PDFs to LaTeX/Markdown.

153Multimodal
AnomalyGPT

AnomalyGPT

Applies large vision-language models to industrial anomaly detection with synthetic data augmentation.

154Data
FaceChain

FaceChain

Alibaba's FaceChain is a personalized portrait generation framework that produces identity-preserving portraits from just a handful of input photos.

155Multimodal
Qwen-VL

Qwen-VL

Alibaba's Qwen-VL is a large-scale vision-language model family with strong performance across captioning, VQA, and visual localization.

156Multimodal
SeamlessM4T

SeamlessM4T

Meta's SeamlessM4T is a unified multilingual and multimodal machine-translation system that handles five translation tasks in one model.

157Multimodal
IT3D

IT3D

Improves Text-to-3D generation by leveraging explicitly synthesized multi-view images in the training loop.

158Multimodal
NeuroImagen

NeuroImagen

Reconstructs visual stimuli images from EEG signals using latent diffusion, opening new windows into visually-evoked brain activity.

159Multimodal
Med-Flamingo

Med-Flamingo

Stanford's Med-Flamingo is a multimodal medical model supporting in-context learning for few-shot medical visual QA.

160Multimodal
OpenFlamingo

OpenFlamingo

An open-source family of autoregressive vision-language models spanning 3B to 9B parameters.

161Multimodal
Dynalang (Agents Model the World with Language)

Dynalang (Agents Model the World with Language)

UC Berkeley's Dynalang agent learns a multimodal world model predicting future text, video, and rewards.

162Agents
Med-PaLM Multimodal

Med-PaLM Multimodal

Introduces a generalist biomedical AI system and a new multimodal biomedical benchmark with 14 tasks.

163Multimodal
Tracking Anything in High Quality

Tracking Anything in High Quality

A framework for high-quality tracking-anything in videos combining segmentation and refinement.

164Multimodal
Foundation Models in Vision

Foundation Models in Vision

A comprehensive survey on foundational models for computer vision and their open research directions.

165Multimodal
WavJourney

WavJourney

Leverages LLMs to orchestrate audio generation models for compositional storytelling.

166Multimodal
Meta-Transformer

Meta-Transformer

A unified framework performing learning across 12 different modalities with a shared backbone.

167Architecture
CM3Leon

CM3Leon

Meta's retrieval-augmented multi-modal language model that generates both text and images.

168Multimodal
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026