🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023Issue 180 · Sep 14 – Sep 20, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,315
Papers
180
Weekly issues
2023
Since
This week · 10 papersView the full issue →
MM1: Multimodal LLM Pre-training

MM1: Multimodal LLM Pre-training

Apple's MM1 paper runs extensive ablations on multimodal LLM pretraining choices and releases a family of models up to 30B parameters that set competitive MLLM pretraining benchmarks.

02Training
Claude 3

Claude 3

Anthropic releases the Claude 3 family (Haiku, Sonnet, Opus), with Opus leapfrogging GPT-4 on many standard benchmarks and bringing frontier multimodal capability plus a much larger context window.

03Evaluation
Robust Evaluation of Reasoning

Robust Evaluation of Reasoning

The paper introduces functional benchmarks that parameterize reasoning problems so the same structural question can be re-instantiated with fresh surface forms, then uses them to expose a large "reasoning gap" in frontier LLMs.

04Reasoning
GaLore

GaLore

GaLore (Gradient Low-Rank Projection) reduces optimizer-state memory during LLM training while still permitting full-parameter updates, unlike LoRA-style adapters that restrict learning to a low-rank subspace.

05Memory
Can LLMs Reason and Plan?

Can LLMs Reason and Plan?

Kambhampati's position paper argues that what looks like reasoning and planning in LLMs is better understood as "universal approximate retrieval" powered by web-scale training.

06Reasoning
RAG for AI-Generated Content

RAG for AI-Generated Content

A survey that extends RAG beyond text, showing how retrieval augmentation is being applied across code, image, audio, video, and 3D generation.

07Retrieval
KnowAgent

KnowAgent

KnowAgent improves LLM-based planning agents by explicitly injecting action knowledge - what the actions are and how they relate - rather than letting the LLM invent its own action space at runtime.

08Agents
Sora Overview

Sora Overview

A comprehensive academic review of OpenAI's Sora, tracing the technical ingredients behind the text-to-video "world simulator" and the opportunities/limitations for the next wave of large vision models.

09Multimodal
SaulLM-7B: LLM for Law

SaulLM-7B: LLM for Law

SaulLM-7B is an open legal-domain LLM built on Mistral 7B and continually pretrained on 30B+ tokens of English legal text, with a companion instruction-tuning recipe.

10Training
Design2Code

Design2Code

Design2Code tackles the front-end engineering problem of turning a visual design into working HTML/CSS and gives the community both a benchmark and strong MLLM baselines.

11Evaluation
TripoSR

TripoSR

TripoSR is a transformer-based single-image 3D reconstruction model that returns a textured mesh in under 0.5 seconds, building on the LRM architecture with a stronger data and training pipeline.

12Training
Genie

Genie

DeepMind's Genie is an 11B-parameter foundation world model trained unsupervised on internet gameplay videos that generates action-controllable 2D worlds from a single image prompt.

13Robotics
Mistral Large

Mistral Large

Mistral AI releases Mistral Large, its flagship closed-weight LLM positioned as the second-ranked API-accessible model behind GPT-4 at launch.

14Agents
The Era of 1-bit LLMs (BitNet b1.58)

The Era of 1-bit LLMs (BitNet b1.58)

Microsoft's BitNet b1.58 shows that restricting every weight to the ternary set {-1, 0, 1} can match full-precision FP16 transformers on perplexity and downstream tasks at the same parameter count.

15Efficiency
Datasets for LLMs: A Comprehensive Survey

Datasets for LLMs: A Comprehensive Survey

A 180+-page survey that catalogs and analyzes the datasets that underpin modern LLM training and evaluation.

16Data
LearnAct

LearnAct

LearnAct lets language agents expand and refine their own action space over time by writing and revising Python functions in response to execution feedback.

17Agents
EMO: Emote Portrait Alive

EMO: Emote Portrait Alive

Alibaba's EMO synthesizes expressive talking-head videos directly from audio, bypassing the intermediate 3D models or facial landmarks used by prior approaches.

18Multimodal
On the Societal Impact of Open Foundation Models

On the Societal Impact of Open Foundation Models

Stanford CRFM's policy paper proposes a rigorous framework for assessing the *marginal* risk of open-weight foundation models relative to closed models and pre-existing technologies.

19Safety
StarCoder 2

StarCoder 2

BigCode releases StarCoder 2, an open family of code LLMs at 3B, 7B, and 15B parameters trained on The Stack v2, a much larger and cleaner code corpus than the original StarCoder.

20Code
LLMs on Tabular Data: A Survey

LLMs on Tabular Data: A Survey

A survey that maps how LLMs are being applied to tabular data tasks - a domain historically dominated by gradient-boosted trees and specialized architectures.

21Evaluation
PlanGPT

PlanGPT

PlanGPT is a domain-specialized LLM framework for urban and spatial planning, built in collaboration with the Chinese Academy of Urban Planning.

22Agents
Stable Diffusion 3

Stable Diffusion 3

Stability AI previews Stable Diffusion 3, a suite of image-generation models from 800M to 8B parameters that shifts to a diffusion-transformer backbone with flow matching.

23Training
Gemma

Gemma

Google DeepMind releases Gemma, a family of open models (2B and 7B) built from the same research stack as Gemini and shipped with both base and instruction-tuned variants.

24Reinforcement Learning
LLMs for Data Annotation

LLMs for Data Annotation

A survey that maps the rapidly growing literature on using LLMs to generate, evaluate, and learn from data annotations.

25Data
180 weeks of AI research · papers per week
Week of Sep 14–20, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026