🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
390 papers · 2023Clear filters →
Voyager

Voyager

An LLM-powered embodied lifelong learning agent in Minecraft exploring autonomously.

313Agents
Gorilla

Gorilla

A fine-tuned LLaMA-based model that surpasses GPT-4 on API call generation.

314Agents
The False Promise of Imitating Proprietary LLMs

The False Promise of Imitating Proprietary LLMs

Berkeley's critical analysis of open-source imitation of proprietary LLMs.

315Training
Sophia

Sophia

A simple, scalable second-order optimizer with negligible per-step overhead.

316Efficiency
The Larger They Are, the Harder They Fail

The Larger They Are, the Harder They Fail

Reveals inverse-scaling failures in LLM code generation.

317Code
Model Evaluation for Extreme Risks

Model Evaluation for Extreme Risks

DeepMind's framework for evaluating models for catastrophic-risk capabilities.

318Evaluation
LLM Research Directions

LLM Research Directions

A list of research directions for students entering LLM research.

319Evaluation
Reinventing RNNs for the Transformer Era (RWKV)

Reinventing RNNs for the Transformer Era (RWKV)

Combines parallelizable training of Transformers with efficient RNN inference.

320Architecture
Drag Your GAN (DragGAN)

Drag Your GAN (DragGAN)

Interactive point-based image manipulation on the generative image manifold.

321Multimodal
Evidence of Meaning in Language Models Trained on Programs

Evidence of Meaning in Language Models Trained on Programs

Argues LMs learn meaning despite only next-token prediction.

322Reasoning
Towards Expert-Level Medical Question Answering (Med-PaLM 2)

Towards Expert-Level Medical Question Answering (Med-PaLM 2)

Google's second-generation medical LLM.

323Evaluation
MEGABYTE

MEGABYTE

Multiscale Transformers for predicting million-byte sequences.

324Architecture
StructGPT

StructGPT

A general framework for LLM reasoning over structured data.

325Reasoning
TinyStories

TinyStories

Explores how small LMs can be and still speak coherent English.

326Data
DoReMi

DoReMi

Optimizes data mixtures for faster language model pretraining.

327Training
CodeT5+

CodeT5+

An open code LLM family for code understanding and generation.

328Code
Symbol tuning

Symbol tuning

Fine-tunes LMs on in-context input-label pairs with natural-language labels replaced by arbitrary symbols.

329Reasoning
Incidental Bilingualism in PaLM's Translation Capability

Incidental Bilingualism in PaLM's Translation Capability

Explores where PaLM's translation ability actually comes from.

330Training
LLM Explains Neurons in LLMs

LLM Explains Neurons in LLMs

OpenAI's automated interpretability pipeline using GPT-4 to explain GPT-2 neurons.

331Safety
PaLM 2

PaLM 2

Google's second-generation PaLM powering Bard and Google products.

332Reasoning
ImageBind

ImageBind

Meta's joint embedding across six modalities at once.

333Multimodal
TidyBot

TidyBot

Combines LLM-based planning and perception with few-shot summarization to infer user preferences.

334Robotics
Unfaithful Explanations in Chain-of-Thought Prompting

Unfaithful Explanations in Chain-of-Thought Prompting

Demonstrates CoT explanations can misrepresent the true reason for a model's prediction.

335Reasoning
InstructBLIP

InstructBLIP

Visual-language instruction tuning built on BLIP-2.

336Multimodal
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026