🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
224 papers · MultimodalClear filters →
KOSMOS-G

KOSMOS-G

Microsoft's KOSMOS-G extends zero-shot image generation to multi-image vision-language input.

169Multimodal
LLaVA-RLHF

LLaVA-RLHF

Adapts factually augmented RLHF to aligning large multimodal models, reducing hallucination without falling into reward-hacking pitfalls.

170Reinforcement Learning
KOSMOS-2.5

KOSMOS-2.5

Microsoft's KOSMOS-2.5 is a multimodal model purpose-built for "machine reading" of text-intensive images.

171Multimodal
ImageBind-LLM

ImageBind-LLM

Shanghai AI Lab's ImageBind-LLM brings six-modality understanding to LLMs via the ImageBind joint embedding space.

172Multimodal
LLaSM (Large Language and Speech Model)

LLaSM (Large Language and Speech Model)

A combined language-and-speech model trained with cross-modal conversational abilities.

173Multimodal
SAM-Med2D

SAM-Med2D

Adapts the Segment Anything Model (SAM) to 2D medical imaging through large-scale medical fine-tuning.

174Multimodal
MVDream

MVDream

ByteDance's MVDream is a multi-view diffusion model that generates geometrically consistent images from multiple viewpoints given a text prompt.

175Multimodal
Nougat

Nougat

Meta's Nougat is a visual transformer for "Neural Optical Understanding for Academic documents" that converts PDFs to LaTeX/Markdown.

176Multimodal
AnomalyGPT

AnomalyGPT

Applies large vision-language models to industrial anomaly detection with synthetic data augmentation.

177Data
FaceChain

FaceChain

Alibaba's FaceChain is a personalized portrait generation framework that produces identity-preserving portraits from just a handful of input photos.

178Multimodal
Qwen-VL

Qwen-VL

Alibaba's Qwen-VL is a large-scale vision-language model family with strong performance across captioning, VQA, and visual localization.

179Multimodal
SeamlessM4T

SeamlessM4T

Meta's SeamlessM4T is a unified multilingual and multimodal machine-translation system that handles five translation tasks in one model.

180Multimodal
IT3D

IT3D

Improves Text-to-3D generation by leveraging explicitly synthesized multi-view images in the training loop.

181Multimodal
NeuroImagen

NeuroImagen

Reconstructs visual stimuli images from EEG signals using latent diffusion, opening new windows into visually-evoked brain activity.

182Multimodal
Med-Flamingo

Med-Flamingo

Stanford's Med-Flamingo is a multimodal medical model supporting in-context learning for few-shot medical visual QA.

183Multimodal
OpenFlamingo

OpenFlamingo

An open-source family of autoregressive vision-language models spanning 3B to 9B parameters.

184Multimodal
Dynalang (Agents Model the World with Language)

Dynalang (Agents Model the World with Language)

UC Berkeley's Dynalang agent learns a multimodal world model predicting future text, video, and rewards.

185Agents
RT-2

RT-2

Google DeepMind's end-to-end vision-language-action model that learns from both web and robotics data to control robots.

186Robotics
Med-PaLM Multimodal

Med-PaLM Multimodal

Introduces a generalist biomedical AI system and a new multimodal biomedical benchmark with 14 tasks.

187Multimodal
Tracking Anything in High Quality

Tracking Anything in High Quality

A framework for high-quality tracking-anything in videos combining segmentation and refinement.

188Multimodal
Foundation Models in Vision

Foundation Models in Vision

A comprehensive survey on foundational models for computer vision and their open research directions.

189Multimodal
WavJourney

WavJourney

Leverages LLMs to orchestrate audio generation models for compositional storytelling.

190Multimodal
Meta-Transformer

Meta-Transformer

A unified framework performing learning across 12 different modalities with a shared backbone.

191Architecture
CM3Leon

CM3Leon

Meta's retrieval-augmented multi-modal language model that generates both text and images.

192Multimodal
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026