🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
420 papers · ReasoningClear filters →
LLMs for Data Annotation

LLMs for Data Annotation

A survey that maps the rapidly growing literature on using LLMs to generate, evaluate, and learn from data annotations.

361Data
Chain-of-Thought Reasoning Without Prompting

Chain-of-Thought Reasoning Without Prompting

DeepMind shows that LLMs often *already* emit chain-of-thought reasoning in alternative decoding paths, and that selecting those paths via confidence lifts reasoning accuracy with no prompt engineering.

362Reasoning
Indirect Reasoning with LLMs (DIR)

Indirect Reasoning with LLMs (DIR)

Direct-Indirect Reasoning augments standard CoT with contrapositive and proof-by-contradiction templates, giving LLMs an explicit way to attack problems they can't solve forward.

363Reasoning
Self-Discover

Self-Discover

Google's Self-Discover lets LLMs compose their own task-specific reasoning strategies from a small library of atomic reasoning modules, at dramatically lower inference cost than self-consistency.

364Reasoning
DeepSeekMath

DeepSeekMath

DeepSeek releases DeepSeekMath 7B, a math-specialized LLM that closes much of the gap to GPT-4 and Gemini-Ultra on MATH by combining better data and a new RL objective.

365Reasoning
LLMs for Mathematical Reasoning

LLMs for Mathematical Reasoning

A survey of the fast-growing literature on using LLMs for mathematical reasoning, from arithmetic word problems to theorem proving.

366Reasoning
AlphaGeometry

AlphaGeometry

DeepMind's AlphaGeometry is a theorem prover that solves Olympiad-level geometry problems at near gold-medallist performance, and crucially, without needing any human demonstrations.

367Reasoning
ReFT (Reinforced Fine-Tuning)

ReFT (Reinforced Fine-Tuning)

ByteDance's ReFT enhances LLM reasoning by combining supervised fine-tuning with online RL that samples alternative reasoning paths, without a learned reward model.

368Reasoning
Chain-of-Table

Chain-of-Table

Google's Chain-of-Table prompts LLMs to iteratively transform a complex table step-by-step to answer questions reliably, extending CoT reasoning to tabular data.

369Reasoning
Generative AI for Math (OpenWebMath / MathPile)

Generative AI for Math (OpenWebMath / MathPile)

Releases a diverse, high-quality math-centric corpus of ~9.5B tokens designed for training math-capable foundation models.

370Reasoning
Survey of Reasoning with Foundation Models

Survey of Reasoning with Foundation Models

A comprehensive survey of reasoning with foundation models, covering tasks, methods, benchmarks, and future directions.

371Reasoning
ReST Meets ReAct

ReST Meets ReAct

Proposes a ReAct-style agent that improves itself via reinforced self-training on its own reasoning traces.

372Agents
Mathematical LLMs Survey

Mathematical LLMs Survey

A survey on the progress of LLMs on mathematical reasoning tasks, covering methods, benchmarks, and open problems.

373Reasoning
Gemini 1.0

Gemini 1.0

Google launches Gemini 1.0, a multimodal family natively designed to reason across text, images, video, audio, and code from the ground up.

374Multimodal
LLMs on Graphs

LLMs on Graphs

A comprehensive overview of the many ways LLMs can be applied to graph-structured data and when each pattern is useful.

375Reasoning
Chain of Code

Chain of Code

DeepMind's Chain of Code extends CoT by encouraging LMs to write pseudocode that mixes real code with LM-simulated sub-routines.

376Reasoning
Medprompt

Medprompt

Microsoft researchers show that careful prompt engineering can push general-purpose GPT-4 to state-of-the-art on medical benchmarks, no domain fine-tuning required.

377Reasoning
System 2 Attention (S2A)

System 2 Attention (S2A)

Meta's S2A uses the LLM's own reasoning to decide what context actually matters, regenerating a clean prompt before the final response step.

378Reasoning
Teaching Small LMs to Reason

Teaching Small LMs to Reason

An approach that teaches smaller language models to explicitly select among reasoning techniques for each problem.

379Reasoning
GPQA

GPQA

A graduate-level Google-proof QA benchmark designed to stress-test reasoning in systems that might exceed human expertise.

380Reasoning
Hitchhiker's Guide From CoT to Agents

Hitchhiker's Guide From CoT to Agents

A survey mapping the conceptual evolution from chain-of-thought reasoning to modern language-agent frameworks.

381Agents
GAIA

GAIA

Meta's GAIA is a benchmark for general AI assistants that requires reasoning, multimodal handling, web browsing, and tool use to solve real-world questions.

382Agents
MedAgents

MedAgents

A collaborative multi-round framework for medical reasoning that uses role-playing LLM agents to improve accuracy and reasoning depth.

383Reasoning
Contrastive Chain-of-Thought

Contrastive Chain-of-Thought

Proposes contrastive CoT prompting where models see both valid *and* invalid reasoning demonstrations to reduce reasoning errors.

384Reasoning
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026