🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023Issue 180 · Sep 14 – Sep 20, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,315
Papers
180
Weekly issues
2023
Since
This week · 10 papersView the full issue →
DeepSeekMath

DeepSeekMath

DeepSeek releases DeepSeekMath 7B, a math-specialized LLM that closes much of the gap to GPT-4 and Gemini-Ultra on MATH by combining better data and a new RL objective.

02Reasoning
LLMs for Table Processing: A Survey

LLMs for Table Processing: A Survey

A survey covering how LLMs and VLMs are used across the full spectrum of table-processing tasks, from classic TableQA to spreadsheet manipulation.

03Evaluation
LLM-based Multi-Agent Systems Survey

LLM-based Multi-Agent Systems Survey

A survey of the fast-growing LLM-based multi-agent systems space, covering both problem-solving applications and "world simulation" research.

04Agents
OLMo

OLMo

Allen AI releases OLMo, a truly open 7B-parameter LLM shipped with training code, pretraining data, full weights, evaluation tooling, and fine-tuning recipes - an answer to the "open-weights but closed-pipeline" releases dominating the space.

05Training
Advances in Multimodal LLMs

Advances in Multimodal LLMs

A comprehensive survey mapping design choices for architecture and training pipeline around multimodal large language models (MLLMs).

06Multimodal
Corrective RAG (CRAG)

Corrective RAG (CRAG)

CRAG adds a self-correcting loop around retrieval so a RAG system can detect and repair bad retrievals instead of feeding them straight into generation.

07Retrieval
LLMs for Mathematical Reasoning

LLMs for Mathematical Reasoning

A survey of the fast-growing literature on using LLMs for mathematical reasoning, from arithmetic word problems to theorem proving.

08Reasoning
Compression Algorithms for LLMs

Compression Algorithms for LLMs

A survey covering the main families of LLM compression techniques and when each one is appropriate.

09Efficiency
MoE-LLaVA

MoE-LLaVA

MoE-LLaVA applies Mixture-of-Experts tuning to the LLaVA vision-language architecture, getting a sparse model with dramatically fewer active parameters at the same compute cost.

10Architecture
Rephrasing the Web (WRAP)

Rephrasing the Web (WRAP)

WRAP uses an off-the-shelf instruction-tuned model to paraphrase web documents into styles like "Wikipedia" or "question-answer format" and trains on the mixture of real + synthetic rephrases.

11Data
The Power of Noise: Redefining Retrieval in RAG

The Power of Noise: Redefining Retrieval in RAG

A study stress-testing the retriever component of RAG systems with surprising results about what actually helps generation.

12Retrieval
Hallucination in LVLMs

Hallucination in LVLMs

A survey specifically scoped to hallucination in Large Vision-Language Models, a phenomenon that differs substantially from text-only LLM hallucination.

13Multimodal
SliceGPT

SliceGPT

Microsoft's SliceGPT is a post-training LLM compression technique that literally slices rows and columns out of weight matrices while preserving zero-shot quality.

14Efficiency
Depth Anything

Depth Anything

A robust monocular depth estimator designed to handle "any image under any circumstance" by scaling self-training on unlabeled data rather than hunting for bigger labeled sets.

15Training
Knowledge Fusion of LLMs (FuseLLM)

Knowledge Fusion of LLMs (FuseLLM)

FuseLLM proposes fusing the capabilities of multiple existing LLMs into a single target model by distilling their output distributions rather than retraining from scratch.

16Training
MambaByte

MambaByte

MambaByte adapts the Mamba state-space architecture to learn directly from raw bytes, bypassing tokenization and all its well-known failure modes.

17Architecture
Diffuse to Choose

Diffuse to Choose

Amazon's Diffuse to Choose is a diffusion-based image-conditioned inpainting model built for "virtual try-on" scenarios where product images must be placed naturally into user scenes.

18Multimodal
WARM (Weighted Averaged Reward Models)

WARM (Weighted Averaged Reward Models)

WARM averages multiple fine-tuned reward models in weight space rather than ensembling their predictions, dramatically reducing RLHF inference cost.

19Reinforcement Learning
Resource-efficient LLMs & Multimodal Foundation Models

Resource-efficient LLMs & Multimodal Foundation Models

A wide-ranging survey of efficiency techniques for LLMs and multimodal foundation models, spanning architecture, algorithms, and system design.

20Efficiency
Red Teaming Visual Language Models

Red Teaming Visual Language Models

Introduces the first dedicated red-teaming benchmark for VLMs, covering vulnerabilities unique to multimodal inputs.

21Evaluation
Lumiere

Lumiere

Google's Lumiere is a space-time diffusion model for text-to-video that generates the entire video duration in a single forward pass rather than cascading short clips.

22Multimodal
Medusa

Medusa

Medusa accelerates LLM inference by bolting on multiple decoding heads that predict several future tokens in parallel, dramatically reducing decoding steps.

23Efficiency
AgentBoard

AgentBoard

AgentBoard is a benchmark and open-source evaluation framework for analytically evaluating LLM agents beyond the usual pass/fail metrics.

24Evaluation
AlphaGeometry

AlphaGeometry

DeepMind's AlphaGeometry is a theorem prover that solves Olympiad-level geometry problems at near gold-medallist performance, and crucially, without needing any human demonstrations.

25Reasoning
180 weeks of AI research · papers per week
Week of Sep 14–20, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026