🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
420 papers · ReasoningClear filters →
Rephrase and Respond (RaR)

Rephrase and Respond (RaR)

An effective prompting method where the LLM rephrases and expands the user's question before answering it.

385Reasoning
EmotionPrompt

EmotionPrompt

Microsoft researchers show that appending emotional stimuli to prompts reliably improves LLM performance across 45 tasks.

386Reasoning
Branch-Solve-Merge (BSM)

Branch-Solve-Merge (BSM)

BSM decomposes LLM tasks into parallel sub-tasks via three LLM-programmed modules: branch, solve, and merge.

387Agents
Llemma

Llemma

Llemma is an open LLM for mathematics built via continued pretraining of Code Llama on the Proof-Pile-2 dataset.

388Data
Hypothesis Search (LLMs Can Learn Rules)

Hypothesis Search (LLMs Can Learn Rules)

A two-stage framework where the LLM learns a rule library for reasoning.

389Reasoning
Meta Chain-of-Thought Prompting (Meta-CoT)

Meta Chain-of-Thought Prompting (Meta-CoT)

A generalizable CoT framework that selects domain-appropriate reasoning patterns for the task at hand.

390Reasoning
Training LLMs with Pause Tokens

Training LLMs with Pause Tokens

CMU shows that adding a learnable `<pause>` token during both pretraining and fine-tuning gives the model extra "thinking time" and improves reasoning.

391Reasoning
Analogical Prompting

Analogical Prompting

Google's Analogical Prompting guides LLM reasoning by having the model self-generate relevant exemplars on the fly.

392Reasoning
Boolformer

Boolformer

The first Transformer trained to perform end-to-end symbolic regression of Boolean functions.

393Reasoning
Logical Chain-of-Thought (LogiCoT)

Logical Chain-of-Thought (LogiCoT)

A neurosymbolic framework that verifies and revises zero-shot CoT reasoning using symbolic-logic principles.

394Reasoning
Chain-of-Verification (CoVe)

Chain-of-Verification (CoVe)

Meta's Chain-of-Verification adds a "deliberation" step where the LLM fact-checks its own draft before finalizing.

395Reasoning
Contrastive Decoding for Reasoning

Contrastive Decoding for Reasoning

Shows that contrastive decoding, a simple inference-time technique, substantially improves reasoning in large LLMs.

396Reasoning
Textbooks Are All You Need II (phi-1.5)

Textbooks Are All You Need II (phi-1.5)

Microsoft's phi-1.5 demonstrates that a 1.3B model trained on "textbook-quality" synthetic data rivals much larger models on reasoning.

397Data
MAmmoTH

MAmmoTH

An open-source LLM family specialized for general mathematical problem solving.

398Reasoning
GPT Solves Math Problems Without a Calculator

GPT Solves Math Problems Without a Calculator

Demonstrates that with sufficient training data, even a small language model can perform accurate multi-digit arithmetic.

399Reasoning
OPRO (LLMs as Optimizers)

OPRO (LLMs as Optimizers)

DeepMind's OPRO uses LLMs as general-purpose optimizers over natural-language-described problems.

400Agents
Graph of Thoughts (GoT)

Graph of Thoughts (GoT)

Generalizes Chain-of-Thought and Tree-of-Thought by modeling LLM reasoning as an arbitrary graph.

401Reasoning
LegalBench

LegalBench

A collaboratively constructed benchmark for measuring legal reasoning in LLMs.

402Evaluation
GPT-4 Code Interpreter for Math

GPT-4 Code Interpreter for Math

A zero-shot prompting technique for GPT-4 Code Interpreter that dramatically boosts math-reasoning accuracy via code self-verification.

403Reasoning
Skeleton-of-Thought (SoT)

Skeleton-of-Thought (SoT)

Microsoft's Skeleton-of-Thought parallelizes LLM generation by first producing an answer skeleton then filling it in concurrently.

404Reasoning
Self-Check

Self-Check

Explores LLM capacity for self-checking on complex reasoning tasks requiring multi-step and non-linear thinking.

405Reasoning
Measuring Faithfulness in Chain-of-Thought Reasoning

Measuring Faithfulness in Chain-of-Thought Reasoning

Anthropic's investigation into whether CoT reasoning actually reflects the model's internal decision process.

406Reasoning
LLMs as General Pattern Machines

LLMs as General Pattern Machines

Demonstrates LLMs serve as general sequence modelers without additional training.

407Reasoning
Teaching Arithmetic to Small Transformers

Teaching Arithmetic to Small Transformers

Trains small transformers on chain-of-thought style data for arithmetic with large gains.

408Reasoning
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026