AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
ReST Meets ReAct
Proposes a ReAct-style agent that improves itself via reinforced self-training on its own reasoning traces.

BabyLLM Challenge Findings
Reports results from a challenge on sample-efficient pretraining using a developmentally plausible corpus.

Weak-to-Strong Generalization
OpenAI's superalignment team shows that weak supervisors can still elicit capabilities from much stronger models - a first empirical signal for scalable oversight.

LLM360
LLM360 is a framework for fully transparent open-source LLM development, with everything from data to training dynamics released.

Gaussian-SLAM
A neural RGBD SLAM method that extends 3D Gaussian Splatting to achieve photorealistic scene reconstruction without sacrificing speed.

Adversarial Diffusion Distillation (SDXL Turbo)
Stability AI's ADD trains a student diffusion model that produces high-quality images in just 1-4 sampling steps.

MEDITRON-70B
EPFL's MEDITRON is an open-source family of medical LLMs at 7B and 70B parameters, continually pretrained on curated medical corpora.

Medprompt
Microsoft researchers show that careful prompt engineering can push general-purpose GPT-4 to state-of-the-art on medical benchmarks, no domain fine-tuning required.

Advancing Long-Context LLMs
A survey of methodologies for improving Transformer long-context capability across pretraining, fine-tuning, and inference stages.

TÜLU 2
Allen AI's TÜLU 2 is a suite of improved open instruction-tuned LLMs and an accompanying study of adaptation best practices.

Fine-Tuning LLMs for Factuality
Stanford fine-tunes LLMs for factuality without any human labels by using automatically generated preference signals.

In-Context Learning Generalization Limits
Investigates whether transformers' in-context learning can generalize beyond the distribution of their pretraining data.

ChipNeMo (LLMs for Chip Design)
NVIDIA's ChipNeMo applies domain-adapted LLMs to industrial chip design workflows.

YaRN (Efficient Context Extension)
YaRN is a compute-efficient method for extending the context window of LLMs well beyond their pretrained length.

FP8-LM
Microsoft's FP8-LM demonstrates that most LLM training variables - gradients, optimizer states - can use FP8 without sacrificing accuracy.

Zephyr
Hugging Face's Zephyr-7B is a 7B parameter LLM whose chat performance rivals much larger chat models aligned with human feedback.

Min-K% Prob (Detecting Pretraining Data)
Proposes Min-K% Prob as an effective detection method for determining whether specific text was in an LLM's pretraining data.

ConvNets Match Vision Transformers
DeepMind shows that strong ConvNet architectures pretrained at scale match ViTs on ImageNet performance at comparable compute.

Llemma
Llemma is an open LLM for mathematics built via continued pretraining of Code Llama on the Proof-Pile-2 dataset.

InstructRetro
NVIDIA introduces Retro 48B, the largest LLM pretrained with retrieval at the time.

FireAct (Language Agent Fine-tuning)
Explores fine-tuning LLMs specifically for language-agent use, demonstrating consistent gains over prompting alone.

RA-DIT (Retrieval-Augmented Dual Instruction Tuning)
Meta's RA-DIT is a lightweight recipe that retrofits LLMs with retrieval capabilities through dual fine-tuning.

The Reversal Curse
Finds that LLMs trained on "A is B" fail to generalize to "B is A" - a surprisingly deep failure of learning.

Graph Neural Prompting (GNP)
A plug-and-play method that injects knowledge-graph information into frozen pretrained LLMs.