AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Rephrase and Respond (RaR)
An effective prompting method where the LLM rephrases and expands the user's question before answering it.

EmotionPrompt
Microsoft researchers show that appending emotional stimuli to prompts reliably improves LLM performance across 45 tasks.

Branch-Solve-Merge (BSM)
BSM decomposes LLM tasks into parallel sub-tasks via three LLM-programmed modules: branch, solve, and merge.

Llemma
Llemma is an open LLM for mathematics built via continued pretraining of Code Llama on the Proof-Pile-2 dataset.

Hypothesis Search (LLMs Can Learn Rules)
A two-stage framework where the LLM learns a rule library for reasoning.

Meta Chain-of-Thought Prompting (Meta-CoT)
A generalizable CoT framework that selects domain-appropriate reasoning patterns for the task at hand.

Training LLMs with Pause Tokens
CMU shows that adding a learnable `<pause>` token during both pretraining and fine-tuning gives the model extra "thinking time" and improves reasoning.

Analogical Prompting
Google's Analogical Prompting guides LLM reasoning by having the model self-generate relevant exemplars on the fly.

Boolformer
The first Transformer trained to perform end-to-end symbolic regression of Boolean functions.

Logical Chain-of-Thought (LogiCoT)
A neurosymbolic framework that verifies and revises zero-shot CoT reasoning using symbolic-logic principles.

Chain-of-Verification (CoVe)
Meta's Chain-of-Verification adds a "deliberation" step where the LLM fact-checks its own draft before finalizing.

Contrastive Decoding for Reasoning
Shows that contrastive decoding, a simple inference-time technique, substantially improves reasoning in large LLMs.

Textbooks Are All You Need II (phi-1.5)
Microsoft's phi-1.5 demonstrates that a 1.3B model trained on "textbook-quality" synthetic data rivals much larger models on reasoning.

MAmmoTH
An open-source LLM family specialized for general mathematical problem solving.

GPT Solves Math Problems Without a Calculator
Demonstrates that with sufficient training data, even a small language model can perform accurate multi-digit arithmetic.

OPRO (LLMs as Optimizers)
DeepMind's OPRO uses LLMs as general-purpose optimizers over natural-language-described problems.

Graph of Thoughts (GoT)
Generalizes Chain-of-Thought and Tree-of-Thought by modeling LLM reasoning as an arbitrary graph.

LegalBench
A collaboratively constructed benchmark for measuring legal reasoning in LLMs.

GPT-4 Code Interpreter for Math
A zero-shot prompting technique for GPT-4 Code Interpreter that dramatically boosts math-reasoning accuracy via code self-verification.

Skeleton-of-Thought (SoT)
Microsoft's Skeleton-of-Thought parallelizes LLM generation by first producing an answer skeleton then filling it in concurrently.

Self-Check
Explores LLM capacity for self-checking on complex reasoning tasks requiring multi-step and non-linear thinking.

Measuring Faithfulness in Chain-of-Thought Reasoning
Anthropic's investigation into whether CoT reasoning actually reflects the model's internal decision process.

LLMs as General Pattern Machines
Demonstrates LLMs serve as general sequence modelers without additional training.

Teaching Arithmetic to Small Transformers
Trains small transformers on chain-of-thought style data for arithmetic with large gains.