AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
LeanDojo
An open-source Lean playground consisting of toolkits, data, models, and benchmarks for theorem proving.

Unifying LLMs & Knowledge Graphs
A roadmap for combining LLMs with knowledge graphs for stronger reasoning.

Imitating Reasoning Process of Larger LLMs (Orca)
Microsoft's 13B model that imitates GPT-4's reasoning traces.

Let's Verify Step by Step
OpenAI's landmark paper on process reward models for mathematical reasoning.

Evidence of Meaning in Language Models Trained on Programs
Argues LMs learn meaning despite only next-token prediction.

Towards Expert-Level Medical Question Answering (Med-PaLM 2)
Google's second-generation medical LLM.

StructGPT
A general framework for LLM reasoning over structured data.

PaLM 2
Google's second-generation PaLM powering Bard and Google products.

Unfaithful Explanations in Chain-of-Thought Prompting
Demonstrates CoT explanations can misrepresent the true reason for a model's prediction.

Learning to Reason and Memorize with Self-Notes
LLMs that deviate from input to explicitly "think" and memorize.

Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
A framework inferring tool sequences for compositional reasoning.

Teaching Large Language Models to Self-Debug
Teaches LLMs to debug their own code via few-shot demonstrations.