🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023Issue 182 · Sep 28 – Oct 4, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
This week · 10 papersView the full issue →
Matryoshka Diffusion Models

Matryoshka Diffusion Models

Apple introduces an end-to-end framework for high-resolution image and video synthesis that denoises across multiple resolutions jointly.

02Multimodal
Spectron

Spectron

Google's Spectron is a spoken-language model trained end-to-end on raw spectrograms rather than text or discrete audio tokens.

03Multimodal
LLMs Meet New Knowledge

LLMs Meet New Knowledge

A benchmark that evaluates how well LLMs handle new knowledge beyond their training cutoff.

04Evaluation
Min-K% Prob (Detecting Pretraining Data)

Min-K% Prob (Detecting Pretraining Data)

Proposes Min-K% Prob as an effective detection method for determining whether specific text was in an LLM's pretraining data.

05Training
ConvNets Match Vision Transformers

ConvNets Match Vision Transformers

DeepMind shows that strong ConvNet architectures pretrained at scale match ViTs on ImageNet performance at comparable compute.

06Training
CommonCanvas

CommonCanvas

Releases CommonCanvas, a text-to-image dataset composed entirely of Creative-Commons-licensed images.

07Data
Managing AI Risks (Bengio, Hinton, et al.)

Managing AI Risks (Bengio, Hinton, et al.)

A high-profile position paper by leading AI researchers laying out risks from upcoming advanced AI systems.

08Safety
Branch-Solve-Merge (BSM)

Branch-Solve-Merge (BSM)

BSM decomposes LLM tasks into parallel sub-tasks via three LLM-programmed modules: branch, solve, and merge.

09Agents
Llemma

Llemma

Llemma is an open LLM for mathematics built via continued pretraining of Code Llama on the Proof-Pile-2 dataset.

10Data
LLMs for Software Engineering

LLMs for Software Engineering

A comprehensive survey of LLMs for software engineering covering models, tasks, evaluation, and open challenges.

11Code
Self-RAG

Self-RAG

Self-RAG trains an LM to adaptively retrieve, generate, and self-critique using special reflection tokens.

12Retrieval
RAG for Long-Form QA

RAG for Long-Form QA

Explores retrieval-augmented LMs specifically on long-form question answering, where RAG failures are more subtle.

13Retrieval
GenBench

GenBench

A Nature Machine Intelligence paper framework for characterizing and understanding generalization research in NLP.

14Evaluation
LLM Self-Explanations

LLM Self-Explanations

Investigates whether LLMs can generate useful feature-attribution explanations for their own outputs.

15Safety
OpenAgents

OpenAgents

An open platform for running and hosting real-world language agents, including three distinct agent types.

16Agents
Eliciting Human Preferences with LLMs

Eliciting Human Preferences with LLMs

Anthropic uses LLMs to guide the task-specification process, eliciting user intent through natural-language dialogue.

17Reinforcement Learning
AutoMix

AutoMix

AutoMix routes queries between LLMs of different sizes based on smaller-model confidence, saving cost without sacrificing quality.

18Efficiency
Video Language Planning

Video Language Planning

Enables synthesizing complex long-horizon video plans for robotics via tree search over vision-language and text-to-video models.

19Multimodal
Ring Attention

Ring Attention

UC Berkeley's Ring Attention scales transformer context to 100M+ tokens by distributing blockwise self-attention across devices in a ring topology.

20Memory
UniSim (Universal Simulator)

UniSim (Universal Simulator)

Google's UniSim learns a universal generative simulator of real-world interactions from diverse video + action data.

21Robotics
Survey on Factuality in LLMs

Survey on Factuality in LLMs

A survey covering evaluation and enhancement techniques for LLM factuality.

22Evaluation
Hypothesis Search (LLMs Can Learn Rules)

Hypothesis Search (LLMs Can Learn Rules)

A two-stage framework where the LLM learns a rule library for reasoning.

23Reasoning
Meta Chain-of-Thought Prompting (Meta-CoT)

Meta Chain-of-Thought Prompting (Meta-CoT)

A generalizable CoT framework that selects domain-appropriate reasoning patterns for the task at hand.

24Reasoning
LLMs for Healthcare Survey

LLMs for Healthcare Survey

A comprehensive overview of LLMs applied to the healthcare domain.

25Evaluation
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026