AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
LLMs in Medicine
A comprehensive survey (300+ papers) of LLMs applied to medicine, from clinical tasks to biomedical research.

Llama Guard
Meta's Llama Guard is a compact, instruction-tuned safety classifier built on Llama 2-7B for input/output moderation in conversational AI.

KTO (Kahneman-Tversky Optimization)
Contextual AI introduces KTO, an alignment objective derived from prospect theory that works with binary "good/bad" signals instead of preference pairs.

Safe Deployment of Generative AI (Nature)
A Nature correspondence arguing that medical professionals - not commercial interests - must drive the development and deployment of generative AI in medicine.

Fine-Tuning LLMs for Factuality
Stanford fine-tunes LLMs for factuality without any human labels by using automatically generated preference signals.

MART (Multi-round Automatic Red-Teaming)
Meta's MART scales LLM safety alignment using fully automatic multi-round red-teaming.

LLMs Can Deceive Users (Trading Agent)
Apollo Research shows that a helpful, honest LLM stock-trading agent can spontaneously deceive users under pressure.

Hallucination in LLMs Survey
A comprehensive survey of hallucination in LLMs, covering taxonomy, causes, evaluation, and mitigation.

Zephyr
Hugging Face's Zephyr-7B is a 7B parameter LLM whose chat performance rivals much larger chat models aligned with human feedback.

Managing AI Risks (Bengio, Hinton, et al.)
A high-profile position paper by leading AI researchers laying out risks from upcoming advanced AI systems.

LLM Self-Explanations
Investigates whether LLMs can generate useful feature-attribution explanations for their own outputs.

LLMs Represent Space and Time
MIT researchers find that LLMs internally encode linear representations of space and time across multiple scales.

LLaVA-RLHF
Adapts factually augmented RLHF to aligning large multimodal models, reducing hallucination without falling into reward-hacking pitfalls.

LLM Alignment Survey
A comprehensive survey of LLM alignment research spanning theoretical foundations to adversarial pressure.

MentaLLaMA
An open-source LLM family specialized for interpretable mental-health analysis on social media.

Rewindable Auto-regressive INference (RAIN)
Shows that unaligned LLMs can produce aligned responses at inference time via self-evaluation and rewinding.

Hallucination Survey (Early)
Classifies hallucination phenomena in LLMs and catalogs evaluation criteria and mitigation strategies.

Explaining Grokking
DeepMind advances our understanding of grokking, predicting and confirming two novel phenomena that test their theory.

Overview of AI Deception
A survey cataloguing empirical examples of AI systems exhibiting deceptive behavior.

LLMs for Illicit Purposes
A survey cataloguing threats and vulnerabilities arising from LLM deployment.

Studying LLM Generalization with Influence Functions
Anthropic scales influence functions to LLMs up to 52B parameters to investigate generalization patterns.

Synthetic Data Reduces Sycophancy
Google shows that fine-tuning on simple synthetic data can significantly reduce LLM sycophancy.

Trustworthy LLMs
Presents a comprehensive framework of categories for assessing LLM trustworthiness.

The Hydra Effect
DeepMind shows that language models exhibit self-repairing behavior when attention heads are ablated.