🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023Issue 182 · Sep 28 – Oct 4, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
This week · 10 papersView the full issue →
SQL-PaLM

SQL-PaLM

An LLM-based Text-to-SQL system built on PaLM-2.

02Code
CodeTF

CodeTF

An open-source Transformer library for state-of-the-art code LLMs.

03Code
QLoRA

QLoRA

Tim Dettmers' breakthrough technique enabling 65B LLM fine-tuning on a single 48GB GPU.

04Training
LIMA

LIMA

Meta's 65B LLaMA fine-tuned on just 1,000 curated examples - showing alignment needs less data than believed.

05Training
Voyager

Voyager

An LLM-powered embodied lifelong learning agent in Minecraft exploring autonomously.

06Agents
Gorilla

Gorilla

A fine-tuned LLaMA-based model that surpasses GPT-4 on API call generation.

07Agents
The False Promise of Imitating Proprietary LLMs

The False Promise of Imitating Proprietary LLMs

Berkeley's critical analysis of open-source imitation of proprietary LLMs.

08Training
Sophia

Sophia

A simple, scalable second-order optimizer with negligible per-step overhead.

09Efficiency
The Larger They Are, the Harder They Fail

The Larger They Are, the Harder They Fail

Reveals inverse-scaling failures in LLM code generation.

10Code
Model Evaluation for Extreme Risks

Model Evaluation for Extreme Risks

DeepMind's framework for evaluating models for catastrophic-risk capabilities.

11Evaluation
LLM Research Directions

LLM Research Directions

A list of research directions for students entering LLM research.

12Evaluation
Reinventing RNNs for the Transformer Era (RWKV)

Reinventing RNNs for the Transformer Era (RWKV)

Combines parallelizable training of Transformers with efficient RNN inference.

13Architecture
Drag Your GAN (DragGAN)

Drag Your GAN (DragGAN)

Interactive point-based image manipulation on the generative image manifold.

14Multimodal
Evidence of Meaning in Language Models Trained on Programs

Evidence of Meaning in Language Models Trained on Programs

Argues LMs learn meaning despite only next-token prediction.

15Reasoning
Towards Expert-Level Medical Question Answering (Med-PaLM 2)

Towards Expert-Level Medical Question Answering (Med-PaLM 2)

Google's second-generation medical LLM.

16Evaluation
MEGABYTE

MEGABYTE

Multiscale Transformers for predicting million-byte sequences.

17Architecture
StructGPT

StructGPT

A general framework for LLM reasoning over structured data.

18Reasoning
TinyStories

TinyStories

Explores how small LMs can be and still speak coherent English.

19Data
DoReMi

DoReMi

Optimizes data mixtures for faster language model pretraining.

20Training
CodeT5+

CodeT5+

An open code LLM family for code understanding and generation.

21Code
Symbol tuning

Symbol tuning

Fine-tunes LMs on in-context input-label pairs with natural-language labels replaced by arbitrary symbols.

22Training
Incidental Bilingualism in PaLM's Translation Capability

Incidental Bilingualism in PaLM's Translation Capability

Explores where PaLM's translation ability actually comes from.

23Training
LLM Explains Neurons in LLMs

LLM Explains Neurons in LLMs

OpenAI's automated interpretability pipeline using GPT-4 to explain GPT-2 neurons.

24Safety
PaLM 2

PaLM 2

Google's second-generation PaLM powering Bard and Google products.

25Reasoning
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026