🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023Issue 180 · Sep 14 – Sep 20, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,315
Papers
180
Weekly issues
2023
Since
This week · 10 papersView the full issue →
Direct Preference Optimization (DPO)

Direct Preference Optimization (DPO)

Rafailov et al.'s simpler alternative to RLHF that rivals full RL-based alignment.

02Reinforcement Learning
SQL-PaLM

SQL-PaLM

An LLM-based Text-to-SQL system built on PaLM-2.

03Code
CodeTF

CodeTF

An open-source Transformer library for state-of-the-art code LLMs.

04Code
QLoRA

QLoRA

Tim Dettmers' breakthrough technique enabling 65B LLM fine-tuning on a single 48GB GPU.

05Training
LIMA

LIMA

Meta's 65B LLaMA fine-tuned on just 1,000 curated examples - showing alignment needs less data than believed.

06Training
Voyager

Voyager

An LLM-powered embodied lifelong learning agent in Minecraft exploring autonomously.

07Agents
Gorilla

Gorilla

A fine-tuned LLaMA-based model that surpasses GPT-4 on API call generation.

08Agents
The False Promise of Imitating Proprietary LLMs

The False Promise of Imitating Proprietary LLMs

Berkeley's critical analysis of open-source imitation of proprietary LLMs.

09Training
Sophia

Sophia

A simple, scalable second-order optimizer with negligible per-step overhead.

10Efficiency
The Larger They Are, the Harder They Fail

The Larger They Are, the Harder They Fail

Reveals inverse-scaling failures in LLM code generation.

11Code
Model Evaluation for Extreme Risks

Model Evaluation for Extreme Risks

DeepMind's framework for evaluating models for catastrophic-risk capabilities.

12Evaluation
LLM Research Directions

LLM Research Directions

A list of research directions for students entering LLM research.

13Evaluation
Reinventing RNNs for the Transformer Era (RWKV)

Reinventing RNNs for the Transformer Era (RWKV)

Combines parallelizable training of Transformers with efficient RNN inference.

14Architecture
Drag Your GAN (DragGAN)

Drag Your GAN (DragGAN)

Interactive point-based image manipulation on the generative image manifold.

15Multimodal
Evidence of Meaning in Language Models Trained on Programs

Evidence of Meaning in Language Models Trained on Programs

Argues LMs learn meaning despite only next-token prediction.

16Reasoning
Towards Expert-Level Medical Question Answering (Med-PaLM 2)

Towards Expert-Level Medical Question Answering (Med-PaLM 2)

Google's second-generation medical LLM.

17Evaluation
MEGABYTE

MEGABYTE

Multiscale Transformers for predicting million-byte sequences.

18Architecture
StructGPT

StructGPT

A general framework for LLM reasoning over structured data.

19Reasoning
TinyStories

TinyStories

Explores how small LMs can be and still speak coherent English.

20Data
DoReMi

DoReMi

Optimizes data mixtures for faster language model pretraining.

21Training
CodeT5+

CodeT5+

An open code LLM family for code understanding and generation.

22Code
Symbol tuning

Symbol tuning

Fine-tunes LMs on in-context input-label pairs with natural-language labels replaced by arbitrary symbols.

23Training
Incidental Bilingualism in PaLM's Translation Capability

Incidental Bilingualism in PaLM's Translation Capability

Explores where PaLM's translation ability actually comes from.

24Training
LLM Explains Neurons in LLMs

LLM Explains Neurons in LLMs

OpenAI's automated interpretability pipeline using GPT-4 to explain GPT-2 neurons.

25Safety
180 weeks of AI research · papers per week
Week of Sep 14–20, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026