🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
390 papers · 2023Clear filters →
In-Context Learning Generalization Limits

In-Context Learning Generalization Limits

Investigates whether transformers' in-context learning can generalize beyond the distribution of their pretraining data.

73Training
MusicGen

MusicGen

Meta's MusicGen is a single-stage transformer LLM for music generation that operates over compressed discrete audio tokens.

74Multimodal
AltUp (Alternating Updates)

AltUp (Alternating Updates)

Google's AltUp lets transformers benefit from wider representations without paying the full compute cost at every layer.

75Architecture
Rephrase and Respond (RaR)

Rephrase and Respond (RaR)

An effective prompting method where the LLM rephrases and expands the user's question before answering it.

76Reasoning
On the Road with GPT-4V

On the Road with GPT-4V

An exhaustive evaluation of GPT-4V applied to autonomous driving scenarios.

77Evaluation
GPT4All Technical Report

GPT4All Technical Report

The GPT4All technical report documents the model family and the open ecosystem built around democratizing local LLMs.

78Data
S-LoRA

S-LoRA

S-LoRA enables serving thousands of LoRA adapters concurrently on a single GPU through memory-paging and custom CUDA kernels.

79Memory
FreshLLMs (FreshQA)

FreshLLMs (FreshQA)

Introduces FreshQA, a dynamic benchmark designed to stress-test LLMs on time-sensitive knowledge.

80Evaluation
MetNet-3

MetNet-3

Google's MetNet-3 is a state-of-the-art neural weather model extending lead time and variable coverage well beyond prior observation-based models.

81Architecture
Evaluating LLMs Survey

Evaluating LLMs Survey

A comprehensive survey of LLM evaluation covering benchmarks, methodologies, and open problems.

82Evaluation
Battle of the Backbones

Battle of the Backbones

A large-scale benchmarking framework that compares vision backbones across a diverse suite of computer vision tasks.

83Architecture
ChipNeMo (LLMs for Chip Design)

ChipNeMo (LLMs for Chip Design)

NVIDIA's ChipNeMo applies domain-adapted LLMs to industrial chip design workflows.

84Training
YaRN (Efficient Context Extension)

YaRN (Efficient Context Extension)

YaRN is a compute-efficient method for extending the context window of LLMs well beyond their pretrained length.

85Training
Open DAC 2023

Open DAC 2023

Meta releases a large DFT dataset for training ML models that predict sorbent-adsorbate interactions in Direct Air Capture (DAC).

86Data
Symmetry in Machine Learning

Symmetry in Machine Learning

A methodological framework for enforcing, discovering, and promoting symmetry in machine learning models.

87Architecture
Next-Generation AlphaFold

Next-Generation AlphaFold

DeepMind previews the next AlphaFold with dramatically expanded scope of biomolecular complexes.

88Architecture
EmotionPrompt

EmotionPrompt

Microsoft researchers show that appending emotional stimuli to prompts reliably improves LLM performance across 45 tasks.

89Reasoning
FP8-LM

FP8-LM

Microsoft's FP8-LM demonstrates that most LLM training variables - gradients, optimizer states - can use FP8 without sacrificing accuracy.

90Efficiency
Zephyr

Zephyr

Hugging Face's Zephyr-7B is a 7B parameter LLM whose chat performance rivals much larger chat models aligned with human feedback.

91Reinforcement Learning
Fact-Checking with LLMs

Fact-Checking with LLMs

Investigates the fact-checking capabilities of frontier LLMs across multiple languages and claim types.

92Retrieval
Matryoshka Diffusion Models

Matryoshka Diffusion Models

Apple introduces an end-to-end framework for high-resolution image and video synthesis that denoises across multiple resolutions jointly.

93Multimodal
Spectron

Spectron

Google's Spectron is a spoken-language model trained end-to-end on raw spectrograms rather than text or discrete audio tokens.

94Multimodal
LLMs Meet New Knowledge

LLMs Meet New Knowledge

A benchmark that evaluates how well LLMs handle new knowledge beyond their training cutoff.

95Evaluation
Min-K% Prob (Detecting Pretraining Data)

Min-K% Prob (Detecting Pretraining Data)

Proposes Min-K% Prob as an effective detection method for determining whether specific text was in an LLM's pretraining data.

96Training
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026