🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023Issue 180 · Sep 14 – Sep 20, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,315
Papers
180
Weekly issues
2023
Since
This week · 10 papersView the full issue →
BabyLLM Challenge Findings

BabyLLM Challenge Findings

Reports results from a challenge on sample-efficient pretraining using a developmentally plausible corpus.

02Training
FunSearch

FunSearch

DeepMind's FunSearch uses LLMs as a mutation operator in an evolutionary loop to discover genuinely new mathematical knowledge.

03Safety
Weak-to-Strong Generalization

Weak-to-Strong Generalization

OpenAI's superalignment team shows that weak supervisors can still elicit capabilities from much stronger models - a first empirical signal for scalable oversight.

04Training
Audiobox

Audiobox

Meta's Audiobox is a unified flow-matching audio model that generates speech, sound effects, and music from natural-language and example prompts.

05Multimodal
Mathematical LLMs Survey

Mathematical LLMs Survey

A survey on the progress of LLMs on mathematical reasoning tasks, covering methods, benchmarks, and open problems.

06Reasoning
LLM360

LLM360

LLM360 is a framework for fully transparent open-source LLM development, with everything from data to training dynamics released.

07Training
LLMs in Medicine

LLMs in Medicine

A comprehensive survey (300+ papers) of LLMs applied to medicine, from clinical tasks to biomedical research.

08Evaluation
Beyond Human Data (ReST-EM)

Beyond Human Data (ReST-EM)

DeepMind's ReST-EM shows that model-generated data plus a reward function can substantially reduce dependence on human-generated data.

09Data
Gaussian-SLAM

Gaussian-SLAM

A neural RGBD SLAM method that extends 3D Gaussian Splatting to achieve photorealistic scene reconstruction without sacrificing speed.

10Training
Pearl

Pearl

Meta's Pearl is a production-ready reinforcement learning agent package designed for real-world deployment constraints.

11Agents
QuIP#

QuIP#

Cornell's QuIP# is a 2-bit LLM quantization scheme that combines lattice codebooks with incoherence processing to close the quality gap to FP16.

12Efficiency
Gemini 1.0

Gemini 1.0

Google launches Gemini 1.0, a multimodal family natively designed to reason across text, images, video, audio, and code from the ground up.

13Multimodal
EfficientSAM

EfficientSAM

Meta's EfficientSAM is a lightweight Segment Anything variant that preserves most of SAM's zero-shot quality at a fraction of the compute.

14Efficiency
Magicoder

Magicoder

Magicoder is a fully open-source code LLM that closes the gap with top commercial code models at only 7B parameters via high-quality synthetic instruction data.

15Code
LLMs on Graphs

LLMs on Graphs

A comprehensive overview of the many ways LLMs can be applied to graph-structured data and when each pattern is useful.

16Reasoning
Llama Guard

Llama Guard

Meta's Llama Guard is a compact, instruction-tuned safety classifier built on Llama 2-7B for input/output moderation in conversational AI.

17Safety
KTO (Kahneman-Tversky Optimization)

KTO (Kahneman-Tversky Optimization)

Contextual AI introduces KTO, an alignment objective derived from prospect theory that works with binary "good/bad" signals instead of preference pairs.

18Reinforcement Learning
Chain of Code

Chain of Code

DeepMind's Chain of Code extends CoT by encouraging LMs to write pseudocode that mixes real code with LM-simulated sub-routines.

19Reasoning
Data Management for LLMs

Data Management for LLMs

A survey of data-management research for LLM pretraining and supervised fine-tuning stages.

20Data
RankZephyr

RankZephyr

RankZephyr is an open-source LLM for listwise zero-shot reranking that bridges the effectiveness gap with GPT-4.

21Retrieval
The Efficiency Spectrum of LLMs

The Efficiency Spectrum of LLMs

A comprehensive review of algorithmic advancements for improving LLM efficiency across the full training-to-inference stack.

22Efficiency
GNoME

GNoME

DeepMind's Graph Networks for Materials Exploration (GNoME) is an AI system that discovered 2.2 million new crystal structures, including 380,000 thermodynamically stable ones.

23Agents
Open-Source LLMs vs. ChatGPT

Open-Source LLMs vs. ChatGPT

A survey cataloguing tasks where open-source LLMs claim to be on par with or better than ChatGPT.

24Evaluation
Adversarial Diffusion Distillation (SDXL Turbo)

Adversarial Diffusion Distillation (SDXL Turbo)

Stability AI's ADD trains a student diffusion model that produces high-quality images in just 1-4 sampling steps.

25Training
180 weeks of AI research · papers per week
Week of Sep 14–20, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026