🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,760
Papers
176
Weekly issues
2023
Since
860 papers · EvaluationClear filters →
CodeT5+

CodeT5+

An open code LLM family for code understanding and generation.

841Code
PaLM 2

PaLM 2

Google's second-generation PaLM powering Bard and Google products.

842Reasoning
InstructBLIP

InstructBLIP

Visual-language instruction tuning built on BLIP-2.

843Multimodal
Active Retrieval Augmented LLMs (FLARE)

Active Retrieval Augmented LLMs (FLARE)

Actively decides when and what to retrieve during generation.

844Retrieval
Are Emergent Abilities of LLMs a Mirage?

Are Emergent Abilities of LLMs a Mirage?

Stanford's critical re-examination of emergent abilities.

845Evaluation
Interpretable ML for Science with PySR

Interpretable ML for Science with PySR

An open-source library for practical symbolic regression in the sciences.

846Safety
PMC-LLaMA

PMC-LLaMA

A LLaMA model fine-tuned on 4.8 million medical papers.

847Training
Distilling Step-by-Step!

Distilling Step-by-Step!

A mechanism to train smaller models that outperform larger LLMs using fewer examples.

848Training
Scaling Transformer to 1M tokens with RMT

Scaling Transformer to 1M tokens with RMT

Recurrent Memory Transformer extends BERT's effective context to 2M tokens.

849Memory
Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond

Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond

A practical guide for practitioners working with LLMs.

850Evaluation
DataComp

DataComp

A multimodal dataset benchmark with 12.8B image-text pairs.

851Multimodal
ChatGPT for Information Extraction

ChatGPT for Information Extraction

A deeper assessment of ChatGPT on information extraction tasks.

852Evaluation
Comparing Physician vs ChatGPT (JAMA)

Comparing Physician vs ChatGPT (JAMA)

A JAMA Internal Medicine study comparing physician and ChatGPT responses.

853Evaluation
Evaluating Verifiability in Generative Search Engines

Evaluating Verifiability in Generative Search Engines

Audits popular generative search engines for citation accuracy.

854Retrieval
OpenAGI: When LLM Meets Domain Experts

OpenAGI: When LLM Meets Domain Experts

An open-source research platform for LLM agents manipulating domain expert models.

855Agents
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

A benchmark using real human standardized exams.

856Evaluation
Segment Everything Everywhere All at Once (SEEM)

Segment Everything Everywhere All at Once (SEEM)

A promptable, interactive segmentation model.

857Multimodal
Eight Things to Know about Large Language Models

Eight Things to Know about Large Language Models

Sam Bowman's influential primer on key LLM considerations.

858Evaluation
A Survey of Large Language Models

A Survey of Large Language Models

A 50-page comprehensive survey on LLMs.

859Training
MACHIAVELLI Benchmark

MACHIAVELLI Benchmark

A benchmark of 134 text-based Choose-Your-Own-Adventure games for measuring ethical trade-offs.

860Evaluation
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026