🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
390 papers · 2023Clear filters →
CM3Leon

CM3Leon

Meta's retrieval-augmented multi-modal language model that generates both text and images.

241Multimodal
Claude 2

Claude 2

Anthropic's second-generation LLM with a detailed model card on safety, alignment, and capabilities.

242Safety
Secrets of RLHF in LLMs

Secrets of RLHF in LLMs

A deep investigation into RLHF with a focus on the inner workings of PPO, including open-source code.

243Reinforcement Learning
LongLLaMA

LongLLaMA

Extends LLaMA's context length using a contrastive training process that reshapes the (key, value) space.

244Memory
Patch n' Pack: NaViT

Patch n' Pack: NaViT

A vision transformer handling any aspect ratio and resolution through sequence packing.

245Architecture
LLMs as General Pattern Machines

LLMs as General Pattern Machines

Demonstrates LLMs serve as general sequence modelers without additional training.

246Reasoning
HyperDreamBooth

HyperDreamBooth

A smaller, faster, and more efficient version of DreamBooth for personalizing text-to-image models.

247Multimodal
Teaching Arithmetic to Small Transformers

Teaching Arithmetic to Small Transformers

Trains small transformers on chain-of-thought style data for arithmetic with large gains.

248Reasoning
AnimateDiff

AnimateDiff

Animates frozen text-to-image diffusion models via a plug-in motion modeling module.

249Multimodal
Generative Pretraining in Multimodality (Emu)

Generative Pretraining in Multimodality (Emu)

A transformer-based multimodal foundation model for generating images and text.

250Multimodal
A Survey on Evaluation of LLMs

A Survey on Evaluation of LLMs

A comprehensive overview of evaluation methods covering what, where, and how to evaluate LLMs.

251Evaluation
How Language Models Use Long Contexts (Lost-in-the-Middle)

How Language Models Use Long Contexts (Lost-in-the-Middle)

Shows LLM performance drops when relevant information is in the middle of a long context.

252Memory
LLMs as Effective Text Rankers

LLMs as Effective Text Rankers

A prompting technique that enables open-source LLMs to perform SOTA text ranking.

253Retrieval
Multimodal Generation with Frozen LLMs

Multimodal Generation with Frozen LLMs

Maps images to LLM token space enabling models like PaLM and GPT-4 to handle visual tasks without parameter updates.

254Multimodal
CodeGen2.5

CodeGen2.5

Salesforce's new 7B code LLM trained on 1.5T tokens and optimized for fast sampling.

255Code
Elastic Decision Transformer

Elastic Decision Transformer

An advance over Decision Transformers that enables trajectory stitching at inference time.

256Reinforcement Learning
Robots That Ask for Help

Robots That Ask for Help

A framework for calibrating LLM-based robot planners so they ask for help when uncertain.

257Robotics
Physics-based Motion Retargeting in Real-Time

Physics-based Motion Retargeting in Real-Time

Uses RL to retarget motions from sparse human sensor data to characters of various morphologies.

258Multimodal
Scaling Transformer to 1 Billion Tokens (LongNet)

Scaling Transformer to 1 Billion Tokens (LongNet)

Microsoft's Transformer variant scaling sequence length past 1B tokens.

259Memory
InterCode

InterCode

A framework treating interactive coding as a reinforcement learning environment.

260Reinforcement Learning
LeanDojo

LeanDojo

An open-source Lean playground consisting of toolkits, data, models, and benchmarks for theorem proving.

261Reasoning
Extending Context Window of LLMs (PI)

Extending Context Window of LLMs (PI)

Position Interpolation extends LLaMA's context to 32K with minimal fine-tuning (within 1000 steps).

262Memory
Computer Vision Through the Lens of Natural Language

Computer Vision Through the Lens of Natural Language

A modular approach solving CV problems by routing through LLM reasoning.

263Multimodal
Visual Navigation Transformer (ViNT)

Visual Navigation Transformer (ViNT)

A foundation model for vision-based robotic navigation built on flexible Transformers.

264Robotics
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026