AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
LLMs in Medicine
A comprehensive survey (300+ papers) of LLMs applied to medicine, from clinical tasks to biomedical research.

Open-Source LLMs vs. ChatGPT
A survey cataloguing tasks where open-source LLMs claim to be on par with or better than ChatGPT.

UniIR
UniIR is a unified instruction-guided multimodal retriever that handles eight retrieval tasks across modalities with a single model.

Mirasol3B
Google's Mirasol3B is a multimodal model that decouples modalities into focused autoregressive components rather than forcing a single fused stream.

GPQA
A graduate-level Google-proof QA benchmark designed to stress-test reasoning in systems that might exceed human expertise.

GAIA
Meta's GAIA is a benchmark for general AI assistants that requires reasoning, multimodal handling, web browsing, and tool use to solve real-world questions.

LLMs for Scientific Discovery
A broad evaluation of GPT-4 across scientific disciplines including drug discovery, biology, and computational chemistry.

MART (Multi-round Automatic Red-Teaming)
Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.

Hallucination in LLMs Survey
A comprehensive survey of hallucination in LLMs, covering taxonomy, causes, evaluation, and mitigation.

On the Road with GPT-4V
An exhaustive evaluation of GPT-4V applied to autonomous driving scenarios.

FreshLLMs (FreshQA)
Introduces FreshQA, a dynamic benchmark designed to stress-test LLMs on time-sensitive knowledge.

Evaluating LLMs Survey
A comprehensive survey of LLM evaluation covering benchmarks, methodologies, and open problems.

Battle of the Backbones
A large-scale benchmarking framework that compares vision backbones across a diverse suite of computer vision tasks.

EmotionPrompt
Microsoft researchers show that appending emotional stimuli to prompts reliably improves LLM performance across 45 tasks.

Fact-Checking with LLMs
Investigates the fact-checking capabilities of frontier LLMs across multiple languages and claim types.

LLMs Meet New Knowledge
A benchmark that evaluates how well LLMs handle new knowledge beyond their training cutoff.

GenBench
A Nature Machine Intelligence paper framework for characterizing and understanding generalization research in NLP.

Survey on Factuality in LLMs
A survey covering evaluation and enhancement techniques for LLM factuality.

LLMs for Healthcare Survey
A comprehensive overview of LLMs applied to the healthcare domain.

LLMs Represent Space and Time
MIT researchers find that LLMs internally encode linear representations of space and time across multiple scales.

Qwen
Alibaba releases the Qwen family of open LLMs with strong tool-use and planning capabilities for language agents.

OWL (LLMs for IT Operations)
Proposes OWL, an LLM specialized for IT operations through self-instruct fine-tuning on IT-specific tasks.

Rewindable Auto-regressive INference (RAIN)
Shows that unaligned LLMs can produce aligned responses at inference time via self-evaluation and rewinding.

Hallucination Survey (Early)
Classifies hallucination phenomena in LLMs and catalogs evaluation criteria and mitigation strategies.