🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
860 papers · EvaluationClear filters →
Long-range Language Modeling with Self-Retrieval

Long-range Language Modeling with Self-Retrieval

Jointly trains a retrieval-augmented LM from scratch for long-range modeling.

817Retrieval
Textbooks Are All You Need (phi-1)

Textbooks Are All You Need (phi-1)

Introduces a 1.3B parameter code LLM trained on textbook-quality data.

818Data
FinGPT

FinGPT

An open-source LLM for the finance sector with a data-centric approach.

819Evaluation
Crowd Workers Widely Use LLMs for Text Production

Crowd Workers Widely Use LLMs for Text Production

Empirical evidence that 33-46% of MTurk crowd workers used LLMs on text tasks.

820Data
Reliability of Watermarks for LLMs

Reliability of Watermarks for LLMs

Studies whether watermarks survive human rewriting and LLM paraphrasing.

821Safety
Benchmarking NN Training Algorithms (AlgoPerf)

Benchmarking NN Training Algorithms (AlgoPerf)

A new benchmark for rigorously evaluating optimizers using realistic workloads.

822Evaluation
Augmenting LLMs with Long-term Memory (LongMem)

Augmenting LLMs with Long-term Memory (LongMem)

Enables LLMs to memorize long history via memory-augmented adaptation.

823Memory
TAPIR

TAPIR

Tracks any queried point on any physical surface throughout a video sequence faster than real-time.

824Evaluation
Mind2Web

Mind2Web

A dataset for evaluating generalist web agents with 2,350 tasks across 137 websites and 31 domains.

825Agents
Tracking Everything Everywhere All at Once (OmniMotion)

Tracking Everything Everywhere All at Once (OmniMotion)

Test-time optimization for dense, long-range motion estimation.

826Multimodal
AlphaDev

AlphaDev

DeepMind's deep RL agent discovering faster sorting algorithms from scratch, now in LLVM.

827Reinforcement Learning
Sparse-Quantized Representation (SpQR)

Sparse-Quantized Representation (SpQR)

Tim Dettmers' near-lossless LLM compression technique.

828Efficiency
MusicGen

MusicGen

A simple and controllable model for music generation using a single-stage Transformer.

829Multimodal
Humor in ChatGPT

Humor in ChatGPT

Explores ChatGPT's capabilities to grasp and reproduce humor.

830Evaluation
Let's Verify Step by Step

Let's Verify Step by Step

OpenAI's landmark paper on process reward models for mathematical reasoning.

831Reasoning
BiomedGPT

BiomedGPT

A unified biomedical GPT for vision, language, and multimodal tasks.

832Multimodal
Thought Cloning

Thought Cloning

Imitation learning framework that learns to think as well as act.

833Agents
MERT

MERT

An acoustic music understanding model with large-scale self-supervised training.

834Multimodal
SQL-PaLM

SQL-PaLM

An LLM-based Text-to-SQL system built on PaLM-2.

835Code
CodeTF

CodeTF

An open-source Transformer library for state-of-the-art code LLMs.

836Code
Model Evaluation for Extreme Risks

Model Evaluation for Extreme Risks

DeepMind's framework for evaluating models for catastrophic-risk capabilities.

837Evaluation
LLM Research Directions

LLM Research Directions

A list of research directions for students entering LLM research.

838Evaluation
Reinventing RNNs for the Transformer Era (RWKV)

Reinventing RNNs for the Transformer Era (RWKV)

Combines parallelizable training of Transformers with efficient RNN inference.

839Architecture
Towards Expert-Level Medical Question Answering (Med-PaLM 2)

Towards Expert-Level Medical Question Answering (Med-PaLM 2)

Google's second-generation medical LLM.

840Evaluation
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026