🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
188 papers · RetrievalClear filters →
FreshLLMs (FreshQA)

FreshLLMs (FreshQA)

Introduces FreshQA, a dynamic benchmark designed to stress-test LLMs on time-sensitive knowledge.

169Evaluation
ChipNeMo (LLMs for Chip Design)

ChipNeMo (LLMs for Chip Design)

NVIDIA's ChipNeMo applies domain-adapted LLMs to industrial chip design workflows.

170Training
Fact-Checking with LLMs

Fact-Checking with LLMs

Investigates the fact-checking capabilities of frontier LLMs across multiple languages and claim types.

171Retrieval
LLMs Meet New Knowledge

LLMs Meet New Knowledge

A benchmark that evaluates how well LLMs handle new knowledge beyond their training cutoff.

172Evaluation
Self-RAG

Self-RAG

Self-RAG trains an LM to adaptively retrieve, generate, and self-critique using special reflection tokens.

173Retrieval
RAG for Long-Form QA

RAG for Long-Form QA

Explores retrieval-augmented LMs specifically on long-form question answering, where RAG failures are more subtle.

174Retrieval
RECOMP (Retrieval-Augmented LMs with Compressors)

RECOMP (Retrieval-Augmented LMs with Compressors)

Proposes two compression approaches to shrink retrieved documents before in-context use.

175Retrieval
InstructRetro

InstructRetro

NVIDIA introduces Retro 48B, the largest LLM pretrained with retrieval at the time.

176Training
Retrieval Meets Long-Context LLMs

Retrieval Meets Long-Context LLMs

NVIDIA's study comparing RAG and long-context LLMs, with the punchline that the two are complementary rather than substitutes.

177Retrieval
RA-DIT (Retrieval-Augmented Dual Instruction Tuning)

RA-DIT (Retrieval-Augmented Dual Instruction Tuning)

Meta's RA-DIT is a lightweight recipe that retrofits LLMs with retrieval capabilities through dual fine-tuning.

178Retrieval
Vector Search with OpenAI Embeddings

Vector Search with OpenAI Embeddings

Argues, via empirical analysis, that dedicated vector databases aren't necessarily required for modern AI-stack search applications.

179Retrieval
FacTool

FacTool

A tool-augmented framework for detecting factual errors in LLM-generated text.

180Retrieval
Retrieve In-Context Examples for LLMs

Retrieve In-Context Examples for LLMs

A framework to iteratively train dense retrievers that identify high-quality in-context examples.

181Retrieval
CM3Leon

CM3Leon

Meta's retrieval-augmented multi-modal language model that generates both text and images.

182Multimodal
LLMs as Effective Text Rankers

LLMs as Effective Text Rankers

A prompting technique that enables open-source LLMs to perform SOTA text ranking.

183Retrieval
LeanDojo

LeanDojo

An open-source Lean playground consisting of toolkits, data, models, and benchmarks for theorem proving.

184Reasoning
Long-range Language Modeling with Self-Retrieval

Long-range Language Modeling with Self-Retrieval

Jointly trains a retrieval-augmented LM from scratch for long-range modeling.

185Retrieval
Unifying LLMs & Knowledge Graphs

Unifying LLMs & Knowledge Graphs

A roadmap for combining LLMs with knowledge graphs for stronger reasoning.

186Retrieval
Active Retrieval Augmented LLMs (FLARE)

Active Retrieval Augmented LLMs (FLARE)

Actively decides when and what to retrieve during generation.

187Retrieval
Evaluating Verifiability in Generative Search Engines

Evaluating Verifiability in Generative Search Engines

Audits popular generative search engines for citation accuracy.

188Retrieval
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026