🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
224 papers · MultimodalClear filters →
A Survey on Retrieval-Augmented Text Generation for LLMs

A Survey on Retrieval-Augmented Text Generation for LLMs

This survey organizes the RAG literature into a four-stage framework (pre-retrieval, retrieval, post-retrieval, generation) and traces the paradigm's evolution alongside open challenges.

121Retrieval
A Survey on State Space Models

A Survey on State Space Models

A comprehensive survey of modern SSMs with a principles-first walkthrough, taxonomy of existing variants, and experimental comparison across NLP, vision, graph, multimodal, point-cloud, event-stream, and time-series tasks.

122Multimodal
OpenEQA

OpenEQA

Meta's OpenEQA is an open-vocabulary benchmark for embodied question answering: 1,600+ human-written questions across 180+ real-world environments, with a calibrated LLM-as-judge metric that tracks human agreement closely.

123Evaluation
Visualization-of-Thought

Visualization-of-Thought

Microsoft's Visualization-of-Thought (VoT) prompts LLMs to emit intermediate "mental images" of their reasoning state, lifting spatial-reasoning accuracy on grid-world tasks and beating multimodal baselines that actually see images.

124Reasoning
SEEDS

SEEDS

Google's Scalable Ensemble Envelope Diffusion Sampler (SEEDS) uses diffusion models to generate very large, physically plausible weather-forecast ensembles conditioned on only one or two operational forecasts.

125Multimodal
Mini-Gemini

Mini-Gemini

Mini-Gemini enhances vision-language models by adding a second high-resolution visual encoder that refines details without increasing the number of visual tokens consumed by the LLM.

126Multimodal
SIMA

SIMA

DeepMind's Scalable Instructable Multiworld Agent (SIMA) is a generalist AI agent that follows natural-language instructions across nine commercial 3D video games like No Man's Sky, Teardown, Valheim, and Space Engineers.

127Agents
MM1: Multimodal LLM Pre-training

MM1: Multimodal LLM Pre-training

Apple's MM1 paper runs extensive ablations on multimodal LLM pretraining choices and releases a family of models up to 30B parameters that set competitive MLLM pretraining benchmarks.

128Training
Claude 3

Claude 3

Anthropic releases the Claude 3 family (Haiku, Sonnet, Opus), with Opus leapfrogging GPT-4 on many standard benchmarks and bringing frontier multimodal capability plus a much larger context window.

129Evaluation
Design2Code

Design2Code

Design2Code tackles the front-end engineering problem of turning a visual design into working HTML/CSS and gives the community both a benchmark and strong MLLM baselines.

130Evaluation
EMO: Emote Portrait Alive

EMO: Emote Portrait Alive

Alibaba's EMO synthesizes expressive talking-head videos directly from audio, bypassing the intermediate 3D models or facial landmarks used by prior approaches.

131Multimodal
LLMs for Data Annotation

LLMs for Data Annotation

A survey that maps the rapidly growing literature on using LLMs to generate, evaluate, and learn from data annotations.

132Data
Sora

Sora

OpenAI unveils Sora, a text-to-video diffusion-transformer that generates coherent, minute-long 1080p videos from natural-language prompts.

133Multimodal
Gemini 1.5

Gemini 1.5

Google DeepMind's Gemini 1.5 is a multimodal MoE LLM that scales context to 1M tokens (10M in research settings) while matching or surpassing Gemini 1.0 Ultra on standard benchmarks.

134Architecture
Large World Model (LWM)

Large World Model (LWM)

UC Berkeley's LWM is an open 7B multimodal model trained on long videos and books that handles context windows up to 1M tokens via RingAttention.

135Memory
LLMs for Table Processing: A Survey

LLMs for Table Processing: A Survey

A survey covering how LLMs and VLMs are used across the full spectrum of table-processing tasks, from classic TableQA to spreadsheet manipulation.

136Evaluation
Advances in Multimodal LLMs

Advances in Multimodal LLMs

A comprehensive survey mapping design choices for architecture and training pipeline around multimodal large language models (MLLMs).

137Multimodal
MoE-LLaVA

MoE-LLaVA

MoE-LLaVA applies Mixture-of-Experts tuning to the LLaVA vision-language architecture, getting a sparse model with dramatically fewer active parameters at the same compute cost.

138Architecture
Hallucination in LVLMs

Hallucination in LVLMs

A survey specifically scoped to hallucination in Large Vision-Language Models, a phenomenon that differs substantially from text-only LLM hallucination.

139Safety
Resource-efficient LLMs & Multimodal Foundation Models

Resource-efficient LLMs & Multimodal Foundation Models

A wide-ranging survey of efficiency techniques for LLMs and multimodal foundation models, spanning architecture, algorithms, and system design.

140Multimodal
Red Teaming Visual Language Models

Red Teaming Visual Language Models

Introduces the first dedicated red-teaming benchmark for VLMs, covering vulnerabilities unique to multimodal inputs.

141Evaluation
Lumiere

Lumiere

Google's Lumiere is a space-time diffusion model for text-to-video that generates the entire video duration in a single forward pass rather than cascading short clips.

142Multimodal
MagicVideo-V2

MagicVideo-V2

ByteDance's MagicVideo-V2 is an end-to-end text-to-video pipeline that stitches together four specialized modules into a high-fidelity generation system.

143Multimodal
Instruct-Imagen

Instruct-Imagen

Google's Instruct-Imagen is a multimodal instruction-tuned image generation model that generalizes across heterogeneous generation tasks, including unseen ones.

144Multimodal
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026