🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
546 papers · EvaluationClear filters →
LLMs in Medicine

LLMs in Medicine

A comprehensive survey (300+ papers) of LLMs applied to medicine, from clinical tasks to biomedical research.

481Evaluation
Open-Source LLMs vs. ChatGPT

Open-Source LLMs vs. ChatGPT

A survey cataloguing tasks where open-source LLMs claim to be on par with or better than ChatGPT.

482Evaluation
UniIR

UniIR

UniIR is a unified instruction-guided multimodal retriever that handles eight retrieval tasks across modalities with a single model.

483Multimodal
Mirasol3B

Mirasol3B

Google's Mirasol3B is a multimodal model that decouples modalities into focused autoregressive components rather than forcing a single fused stream.

484Multimodal
GPQA

GPQA

A graduate-level Google-proof QA benchmark designed to stress-test reasoning in systems that might exceed human expertise.

485Reasoning
GAIA

GAIA

Meta's GAIA is a benchmark for general AI assistants that requires reasoning, multimodal handling, web browsing, and tool use to solve real-world questions.

486Agents
LLMs for Scientific Discovery

LLMs for Scientific Discovery

A broad evaluation of GPT-4 across scientific disciplines including drug discovery, biology, and computational chemistry.

487Evaluation
MART (Multi-round Automatic Red-Teaming)

MART (Multi-round Automatic Red-Teaming)

Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.

488Safety
Hallucination in LLMs Survey

Hallucination in LLMs Survey

A comprehensive survey of hallucination in LLMs, covering taxonomy, causes, evaluation, and mitigation.

489Safety
On the Road with GPT-4V

On the Road with GPT-4V

An exhaustive evaluation of GPT-4V applied to autonomous driving scenarios.

490Evaluation
FreshLLMs (FreshQA)

FreshLLMs (FreshQA)

Introduces FreshQA, a dynamic benchmark designed to stress-test LLMs on time-sensitive knowledge.

491Evaluation
Evaluating LLMs Survey

Evaluating LLMs Survey

A comprehensive survey of LLM evaluation covering benchmarks, methodologies, and open problems.

492Evaluation
Battle of the Backbones

Battle of the Backbones

A large-scale benchmarking framework that compares vision backbones across a diverse suite of computer vision tasks.

493Architecture
EmotionPrompt

EmotionPrompt

Microsoft researchers show that appending emotional stimuli to prompts reliably improves LLM performance across 45 tasks.

494Reasoning
Fact-Checking with LLMs

Fact-Checking with LLMs

Investigates the fact-checking capabilities of frontier LLMs across multiple languages and claim types.

495Retrieval
LLMs Meet New Knowledge

LLMs Meet New Knowledge

A benchmark that evaluates how well LLMs handle new knowledge beyond their training cutoff.

496Evaluation
GenBench

GenBench

A Nature Machine Intelligence paper framework for characterizing and understanding generalization research in NLP.

497Evaluation
Survey on Factuality in LLMs

Survey on Factuality in LLMs

A survey covering evaluation and enhancement techniques for LLM factuality.

498Evaluation
LLMs for Healthcare Survey

LLMs for Healthcare Survey

A comprehensive overview of LLMs applied to the healthcare domain.

499Evaluation
LLMs Represent Space and Time

LLMs Represent Space and Time

MIT researchers find that LLMs internally encode linear representations of space and time across multiple scales.

500Safety
Qwen

Qwen

Alibaba releases the Qwen family of open LLMs with strong tool-use and planning capabilities for language agents.

501Agents
OWL (LLMs for IT Operations)

OWL (LLMs for IT Operations)

Proposes OWL, an LLM specialized for IT operations through self-instruct fine-tuning on IT-specific tasks.

502Evaluation
Rewindable Auto-regressive INference (RAIN)

Rewindable Auto-regressive INference (RAIN)

Shows that unaligned LLMs can produce aligned responses at inference time via self-evaluation and rewinding.

503Safety
Hallucination Survey (Early)

Hallucination Survey (Early)

Classifies hallucination phenomena in LLMs and catalogs evaluation criteria and mitigation strategies.

504Safety
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026