🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,315
Papers
180
Weekly issues
2023
Since
188 papers · EfficiencyClear filters →
RECOMP (Retrieval-Augmented LMs with Compressors)

RECOMP (Retrieval-Augmented LMs with Compressors)

Proposes two compression approaches to shrink retrieved documents before in-context use.

169Retrieval
Language Modeling Is Compression

Language Modeling Is Compression

DeepMind empirically revisits the theoretical equivalence between prediction and compression, applied to modern LLMs.

170Efficiency
Model Compression for LLMs Survey

Model Compression for LLMs Survey

A survey of recent model-compression techniques applied specifically to LLMs.

171Efficiency
Outlines (Efficient Guided Generation)

Outlines (Efficient Guided Generation)

A library for guided LLM text generation that enforces structural constraints with minimal overhead.

172Efficiency
SynJax

SynJax

DeepMind's SynJax is a JAX-based library for efficient vectorized inference in structured distributions.

173Efficiency
Skeleton-of-Thought (SoT)

Skeleton-of-Thought (SoT)

Microsoft's Skeleton-of-Thought parallelizes LLM generation by first producing an answer skeleton then filling it in concurrently.

174Reasoning
FlashAttention-2

FlashAttention-2

Tri Dao's follow-up to FlashAttention, dramatically improving attention throughput on modern GPUs.

175Efficiency
Retentive Network (RetNet)

Retentive Network (RetNet)

Microsoft's proposed foundation architecture aiming to replace Transformer attention for LLMs.

176Architecture
Physics-based Motion Retargeting in Real-Time

Physics-based Motion Retargeting in Real-Time

Uses RL to retarget motions from sparse human sensor data to characters of various morphologies.

177Multimodal
LOMO

LOMO

A memory-efficient optimizer that combines gradient computation and parameter update in one step.

178Efficiency
MotionGPT

MotionGPT

Generates consecutive human motions from multimodal control signals via LLM instructions.

179Multimodal
Wanda

Wanda

A simple, effective pruning approach for LLMs requiring no retraining.

180Efficiency
Sparse-Quantized Representation (SpQR)

Sparse-Quantized Representation (SpQR)

Tim Dettmers' near-lossless LLM compression technique.

181Efficiency
Fine-Tuning Language Models with Just Forward Passes (MeZO)

Fine-Tuning Language Models with Just Forward Passes (MeZO)

A memory-efficient zeroth-order optimizer for LLM fine-tuning.

182Training
QLoRA

QLoRA

Tim Dettmers' breakthrough technique enabling 65B LLM fine-tuning on a single 48GB GPU.

183Training
Sophia

Sophia

A simple, scalable second-order optimizer with negligible per-step overhead.

184Efficiency
Reinventing RNNs for the Transformer Era (RWKV)

Reinventing RNNs for the Transformer Era (RWKV)

Combines parallelizable training of Transformers with efficient RNN inference.

185Architecture
MEGABYTE

MEGABYTE

Multiscale Transformers for predicting million-byte sequences.

186Architecture
FrugalGPT

FrugalGPT

Strategies to reduce LLM inference cost while improving performance.

187Efficiency
Learning to Compress Prompts with Gist Tokens

Learning to Compress Prompts with Gist Tokens

Trains LMs to compress prompts into reusable "gist" tokens.

188Efficiency
180 weeks of AI research · papers per week
Week of Sep 14–20, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026