🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
392 papers · TrainingClear filters →
LOMO

LOMO

A memory-efficient optimizer that combines gradient computation and parameter update in one step.

361Efficiency
LMFlow

LMFlow

An extensible and lightweight toolkit for fine-tuning and inference of large foundation models.

362Training
Fine-Tuning Language Models with Just Forward Passes (MeZO)

Fine-Tuning Language Models with Just Forward Passes (MeZO)

A memory-efficient zeroth-order optimizer for LLM fine-tuning.

363Training
MERT

MERT

An acoustic music understanding model with large-scale self-supervised training.

364Multimodal
Bytes Are All You Need

Bytes Are All You Need

Performs classification directly on file bytes without decoding.

365Training
QLoRA

QLoRA

Tim Dettmers' breakthrough technique enabling 65B LLM fine-tuning on a single 48GB GPU.

366Training
LIMA

LIMA

Meta's 65B LLaMA fine-tuned on just 1,000 curated examples - showing alignment needs less data than believed.

367Training
The False Promise of Imitating Proprietary LLMs

The False Promise of Imitating Proprietary LLMs

Berkeley's critical analysis of open-source imitation of proprietary LLMs.

368Training
Sophia

Sophia

A simple, scalable second-order optimizer with negligible per-step overhead.

369Efficiency
DoReMi

DoReMi

Optimizes data mixtures for faster language model pretraining.

370Training
CodeT5+

CodeT5+

An open code LLM family for code understanding and generation.

371Code
Symbol tuning

Symbol tuning

Fine-tunes LMs on in-context input-label pairs with natural-language labels replaced by arbitrary symbols.

372Training
Incidental Bilingualism in PaLM's Translation Capability

Incidental Bilingualism in PaLM's Translation Capability

Explores where PaLM's translation ability actually comes from.

373Training
InstructBLIP

InstructBLIP

Visual-language instruction tuning built on BLIP-2.

374Multimodal
scGPT

scGPT

A foundation model for single-cell multi-omics pretrained on 10 million cells.

375Training
PMC-LLaMA

PMC-LLaMA

A LLaMA model fine-tuned on 4.8 million medical papers.

376Training
Distilling Step-by-Step!

Distilling Step-by-Step!

A mechanism to train smaller models that outperform larger LLMs using fewer examples.

377Training
Poisoning Language Models During Instruction Tuning

Poisoning Language Models During Instruction Tuning

Shows adversaries can poison LLMs via instruction tuning data.

378Training
A Cookbook of Self-Supervised Learning

A Cookbook of Self-Supervised Learning

A comprehensive overview of SSL techniques and practical considerations.

379Training
Stable and Low-Precision Training for Large-Scale Vision-Language Models

Stable and Low-Precision Training for Large-Scale Vision-Language Models

Methods for accelerating and stabilizing large VLM training.

380Training
DINOv2

DINOv2

Meta's self-supervised vision foundation model producing robust features without labels.

381Training
Visual Instruction Tuning (LLaVA)

Visual Instruction Tuning (LLaVA)

Uses language-only GPT-4 to generate multimodal instruction-following data.

382Multimodal
ChatGPT: Applications, Opportunities, and Threats

ChatGPT: Applications, Opportunities, and Threats

A comprehensive overview of ChatGPT's applications and risks.

383Training
Zip-NeRF: Anti-Aliased Grid-Based Neural Radiance Fields

Zip-NeRF: Anti-Aliased Grid-Based Neural Radiance Fields

Combines mip-NeRF 360 with grid-based models for 22x faster training.

384Training
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026