AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
FreshLLMs (FreshQA)
Introduces FreshQA, a dynamic benchmark designed to stress-test LLMs on time-sensitive knowledge.

ChipNeMo (LLMs for Chip Design)
NVIDIA's ChipNeMo applies domain-adapted LLMs to industrial chip design workflows.

Fact-Checking with LLMs
Investigates the fact-checking capabilities of frontier LLMs across multiple languages and claim types.

LLMs Meet New Knowledge
A benchmark that evaluates how well LLMs handle new knowledge beyond their training cutoff.

Self-RAG
Self-RAG trains an LM to adaptively retrieve, generate, and self-critique using special reflection tokens.

RAG for Long-Form QA
Explores retrieval-augmented LMs specifically on long-form question answering, where RAG failures are more subtle.

RECOMP (Retrieval-Augmented LMs with Compressors)
Proposes two compression approaches to shrink retrieved documents before in-context use.

InstructRetro
NVIDIA introduces Retro 48B, the largest LLM pretrained with retrieval at the time.

Retrieval Meets Long-Context LLMs
NVIDIA's study comparing RAG and long-context LLMs, with the punchline that the two are complementary rather than substitutes.

RA-DIT (Retrieval-Augmented Dual Instruction Tuning)
Meta's RA-DIT is a lightweight recipe that retrofits LLMs with retrieval capabilities through dual fine-tuning.

Vector Search with OpenAI Embeddings
Argues, via empirical analysis, that dedicated vector databases aren't necessarily required for modern AI-stack search applications.

FacTool
A tool-augmented framework for detecting factual errors in LLM-generated text.

Retrieve In-Context Examples for LLMs
A framework to iteratively train dense retrievers that identify high-quality in-context examples.

CM3Leon
Meta's retrieval-augmented multi-modal language model that generates both text and images.

LLMs as Effective Text Rankers
A prompting technique that enables open-source LLMs to perform SOTA text ranking.

LeanDojo
An open-source Lean playground consisting of toolkits, data, models, and benchmarks for theorem proving.

Long-range Language Modeling with Self-Retrieval
Jointly trains a retrieval-augmented LM from scratch for long-range modeling.

Unifying LLMs & Knowledge Graphs
A roadmap for combining LLMs with knowledge graphs for stronger reasoning.

Active Retrieval Augmented LLMs (FLARE)
Actively decides when and what to retrieve during generation.

Evaluating Verifiability in Generative Search Engines
Audits popular generative search engines for citation accuracy.