🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
188 papers · RetrievalClear filters →
How Faithful are RAG Models? (ClashEval)

How Faithful are RAG Models? (ClashEval)

ClashEval constructs a 1,200-question benchmark across six domains with intentionally corrupted retrieved documents to measure when RAG helps and when it misleads GPT-4 and other top LLMs.

145Retrieval
A Survey on Retrieval-Augmented Text Generation for LLMs

A Survey on Retrieval-Augmented Text Generation for LLMs

This survey organizes the RAG literature into a four-stage framework (pre-retrieval, retrieval, post-retrieval, generation) and traces the paradigm's evolution alongside open challenges.

146Retrieval
Reducing Hallucination in Structured Outputs via RAG

Reducing Hallucination in Structured Outputs via RAG

This paper deploys a compact RAG pipeline - small retriever plus small LM - for an enterprise workflow-generation task and shows it reduces hallucination while improving out-of-domain generalization vs a baseline LLM.

147Retrieval
The Influence Between NLP and Other Fields

The Influence Between NLP and Other Fields

This EMNLP 2023 analysis quantifies NLP's cross-disciplinary engagement using a Citation Field Diversity Index across 23 academic fields. The headline: NLP has become dramatically more insular over four decades.

148Retrieval
FollowIR

FollowIR

FollowIR is both a benchmark and a training set for teaching retrieval models to follow real-world, instruction-style queries rather than just match keywords.

149Evaluation
TacticAI

TacticAI

Google DeepMind, in collaboration with Liverpool FC, releases TacticAI, a geometric deep-learning system that analyzes football corner kicks and suggests alternative tactics for coaches to explore.

150Retrieval
RAFT: Retrieval-Augmented Fine-Tuning

RAFT: Retrieval-Augmented Fine-Tuning

RAFT is a fine-tuning recipe that teaches LLMs to handle distractor documents during RAG and to answer with CoT-style citations to retrieved passages.

151Retrieval
Retrieval Augmented Thoughts (RAT)

Retrieval Augmented Thoughts (RAT)

RAT augments chain-of-thought by iteratively rewriting each reasoning step using retrieved context, sharply reducing hallucination on long-horizon generation tasks.

152Retrieval
Knowledge Conflicts for LLMs

Knowledge Conflicts for LLMs

A survey that maps the landscape of knowledge conflicts in LLMs, covering how they arise, how models behave under them, and how to mitigate them.

153Retrieval
C4AI Command-R

C4AI Command-R

Cohere for AI releases Command-R, a 35B open-weight LLM tuned specifically for retrieval-augmented generation, tool use, and multilingual workflows.

154Retrieval
RAG for AI-Generated Content

RAG for AI-Generated Content

A survey that extends RAG beyond text, showing how retrieval augmentation is being applied across code, image, audio, video, and 3D generation.

155Retrieval
GRIT

GRIT

GRIT (Generative Representational Instruction Tuning) trains a single LLM to handle both generative and embedding tasks, switching behavior based on instructions.

156Retrieval
AnyTool

AnyTool

AnyTool is a training-free LLM agent that scales tool-use to 16K+ Rapid APIs through a hierarchical retriever and a self-reflective solver.

157Agents
Corrective RAG (CRAG)

Corrective RAG (CRAG)

CRAG adds a self-correcting loop around retrieval so a RAG system can detect and repair bad retrievals instead of feeding them straight into generation.

158Retrieval
The Power of Noise: Redefining Retrieval in RAG

The Power of Noise: Redefining Retrieval in RAG

A study stress-testing the retriever component of RAG systems with surprising results about what actually helps generation.

159Retrieval
RAG vs. Finetuning

RAG vs. Finetuning

Microsoft researchers systematically compare RAG and fine-tuning (and their combination) on LLMs like Llama 2 and GPT-4 using an agricultural domain dataset.

160Retrieval
Mitigating Hallucination in LLMs

Mitigating Hallucination in LLMs

A survey cataloging 32 hallucination-mitigation techniques and organizing them into a practical taxonomy.

161Safety
Exploiting Novel GPT-4 APIs

Exploiting Novel GPT-4 APIs

A red-team study of three newer GPT-4 API surfaces - fine-tuning, function calling, and knowledge retrieval - that reveals each introduces new attack vectors.

162Training
LLaRA

LLaRA

LLaRA adapts a decoder-only LLM for dense retrieval via two tailored pretext tasks that leverage text embeddings from the LLM itself.

163Retrieval
RAG for LLMs

RAG for LLMs

A broad survey of Retrieval-Augmented Generation research, organizing the rapidly growing literature into a coherent map.

164Retrieval
RankZephyr

RankZephyr

RankZephyr is an open-source LLM for listwise zero-shot reranking that bridges the effectiveness gap with GPT-4.

165Retrieval
UniIR

UniIR

UniIR is a unified instruction-guided multimodal retriever that handles eight retrieval tasks across modalities with a single model.

166Multimodal
Chain-of-Note (CoN)

Chain-of-Note (CoN)

Tencent's Chain-of-Note adds an explicit note-taking step to RAG so the model can evaluate retrieved evidence before answering.

167Retrieval
Learning to Filter Context for RAG (FILCO)

Learning to Filter Context for RAG (FILCO)

CMU's FILCO improves RAG by training a dedicated model to filter retrieved contexts before they reach the generator.

168Retrieval
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026