🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

1,761
Papers
176
Weekly issues
2023
Since
860 papers · EvaluationClear filters →
LLMs for Scientific Discovery

LLMs for Scientific Discovery

A broad evaluation of GPT-4 across scientific disciplines including drug discovery, biology, and computational chemistry.

721Evaluation
Contrastive Chain-of-Thought

Contrastive Chain-of-Thought

Proposes contrastive CoT prompting where models see both valid *and* invalid reasoning demonstrations to reduce reasoning errors.

722Reasoning
Survey on Language Models for Code

Survey on Language Models for Code

A comprehensive survey of LLMs for code covering 50+ models, 30+ evaluation tasks, and 500 related works.

723Evaluation
Learning to Filter Context for RAG (FILCO)

Learning to Filter Context for RAG (FILCO)

CMU's FILCO improves RAG by training a dedicated model to filter retrieved contexts before they reach the generator.

724Retrieval
MART (Multi-round Automatic Red-Teaming)

MART (Multi-round Automatic Red-Teaming)

Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.

725Safety
Hallucination in LLMs Survey

Hallucination in LLMs Survey

A comprehensive survey of hallucination in LLMs, covering taxonomy, causes, evaluation, and mitigation.

726Safety
In-Context Learning Generalization Limits

In-Context Learning Generalization Limits

Investigates whether transformers' in-context learning can generalize beyond the distribution of their pretraining data.

727Training
Rephrase and Respond (RaR)

Rephrase and Respond (RaR)

An effective prompting method where the LLM rephrases and expands the user's question before answering it.

728Reasoning
On the Road with GPT-4V

On the Road with GPT-4V

An exhaustive evaluation of GPT-4V applied to autonomous driving scenarios.

729Evaluation
FreshLLMs (FreshQA)

FreshLLMs (FreshQA)

Introduces FreshQA, a dynamic benchmark designed to stress-test LLMs on time-sensitive knowledge.

730Evaluation
Evaluating LLMs Survey

Evaluating LLMs Survey

A comprehensive survey of LLM evaluation covering benchmarks, methodologies, and open problems.

731Evaluation
Battle of the Backbones

Battle of the Backbones

A large-scale benchmarking framework that compares vision backbones across a diverse suite of computer vision tasks.

732Architecture
ChipNeMo (LLMs for Chip Design)

ChipNeMo (LLMs for Chip Design)

NVIDIA's ChipNeMo applies domain-adapted LLMs to industrial chip design workflows.

733Training
EmotionPrompt

EmotionPrompt

Microsoft researchers show that appending emotional stimuli to prompts reliably improves LLM performance across 45 tasks.

734Reasoning
Zephyr

Zephyr

Hugging Face's Zephyr-7B is a 7B parameter LLM whose chat performance rivals much larger chat models aligned with human feedback.

735Reinforcement Learning
Fact-Checking with LLMs

Fact-Checking with LLMs

Investigates the fact-checking capabilities of frontier LLMs across multiple languages and claim types.

736Retrieval
LLMs Meet New Knowledge

LLMs Meet New Knowledge

A benchmark that evaluates how well LLMs handle new knowledge beyond their training cutoff.

737Evaluation
Min-K% Prob (Detecting Pretraining Data)

Min-K% Prob (Detecting Pretraining Data)

Proposes Min-K% Prob as an effective detection method for determining whether specific text was in an LLM's pretraining data.

738Training
Branch-Solve-Merge (BSM)

Branch-Solve-Merge (BSM)

BSM decomposes LLM tasks into parallel sub-tasks via three LLM-programmed modules: branch, solve, and merge.

739Agents
Llemma

Llemma

Llemma is an open LLM for mathematics built via continued pretraining of Code Llama on the Proof-Pile-2 dataset.

740Data
LLMs for Software Engineering

LLMs for Software Engineering

A comprehensive survey of LLMs for software engineering covering models, tasks, evaluation, and open challenges.

741Code
Self-RAG

Self-RAG

Self-RAG trains an LM to adaptively retrieve, generate, and self-critique using special reflection tokens.

742Retrieval
RAG for Long-Form QA

RAG for Long-Form QA

Explores retrieval-augmented LMs specifically on long-form question answering, where RAG failures are more subtle.

743Retrieval
GenBench

GenBench

A Nature Machine Intelligence paper framework for characterizing and understanding generalization research in NLP.

744Evaluation
176 weeks of AI research · papers per week
Week of Aug 17–23, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Apr 2023Hover a week to inspect · select to openAug 2026