🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
420 papers · ReasoningClear filters →
Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit Signal

Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit Signal

Yunxiang Mo, Donghao Zhao (HKUST) and Hejia Geng (University of Oxford) preregister a sweep of 3,520 self-consensus early-exit rules and find that none clears three acceptance gates, because agreement measures answer persistence rather than reasoning termination.

49Reasoning
TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards

TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards

Rui Sun, Zhan Shi and Bing He (independent researchers) train diagnostic reasoning agents by sampling an intervention, injecting it into a simulator, and generating the observations it would produce, so the hidden intervention supplies an oracle label for a task where real ground truth would require expert investigation.

50Reinforcement Learning
Prompt Repetition Improves Non-Reasoning LLMs

Prompt Repetition Improves Non-Reasoning LLMs

Yaniv Leviathan, Matan Kalman and Yossi Matias at Google Research report that simply repeating the input prompt improves non-reasoning model performance across Gemini, GPT, Claude and DeepSeek.

51Reasoning
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs

Seogyeong Jeong and colleagues at KAIST and NAVER AI Lab test whether the functional operations inside a chain of thought, such as problem formulation, goal decomposition and deduction, have distinct geometric structure in hidden representations.

52Reasoning
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Ji Soo Lee and colleagues at Meta and KAIST build WearableQA from the wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements each.

53Evaluation
LOCI: A Locator-Critic with Refinement Loop

LOCI: A Locator-Critic with Refinement Loop

Walid Bousselham, Mathilde Caron, Arsha Nagrani and Cordelia Schmid at Google DeepMind argue that VLM failures on hard visual tasks come from failing to locate the relevant detail, not from weak high-level reasoning, and fix it with a two-agent loop that needs no training.

54Agents
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Jacqueline He and colleagues at Meta AI, the University of Washington and Princeton show that standard knowledge distillation helps reasoning and hurts factual recall during mid-training, and trace the cause to teacher confidence.

55Training
What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation

What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation

Daisuke Kikuta (NTT) studies revision propagation, where a user asks for one local change and the model must find and update every dependent part of an artifact whose dependencies are buried in the conversation history.

56Reasoning
LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

Sijie Wang, Zhiqiang Tan, Xinrui Yang and Shaohuai Shi at Harbin Institute of Technology Shenzhen remove the recomputation step that DanceGRPO and FlowGRPO perform after rollout, which is mathematically redundant when rollout and update share a backend under on-policy training.

57Reasoning
RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory

RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory

Yuxiang Wang and colleagues fix two limitations of looped-layer latent recurrence at once, letting each iteration attend to its own earlier states and letting the model decide how many loops a given input deserves.

58Memory
It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

Yigit Utku Bulut supplies the two counterfactual controls that the breakthrough-moment and early-legible-fate readings of reasoning traces have been missing, and both readings largely fail to survive them.

59Reasoning
GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving

GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving

Qiankun Ma and colleagues point out that every KV compression method fixes the per-request budget in advance and only decides what to keep, then make capacity itself a runtime resource that grows on demand.

60Efficiency
Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

Leqi Zheng and colleagues propose Gradient-Aligned Reward, which builds a dense reasoning-aware reward by comparing each rollout's gradient direction to an expert-anchor gradient, using expert solutions already sitting in the training corpus.

61Safety
Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

Kevin Du, Alexander Hoyle, Laura Ruis and Acyr Locatelli (ETH Zurich, work done during an internship at Cohere) test whether the text of a reasoning step actually encodes how much that step mattered, using Monte Carlo advantage as ground truth, and find only partial recoverability.

62Reasoning
RuleMem: Active Rule Memory for Long-Term Conversational Agents

RuleMem: Active Rule Memory for Long-Term Conversational Agents

Xingyuan Zeng and colleagues propose RuleMem, which induces reusable natural-language Horn clauses from conversation history so that agent memory actively guides retrieval and reasoning instead of sitting as passively stored facts.

63Memory
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

Heng Wang and colleagues at Salesforce AI Research, UIUC and Cornell show that the token-importance scores every KV cache evictor computes are close to worthless: evicting uniformly at random inside each head matches the strongest prior evictor while serving 32 to 43% higher throughput.

64Memory
Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers

Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers

Where you put a reasoning trace changes long-context accuracy by up to 50 points. Transformers process causally, so a task state discovered late cannot guide reading that already happened, and Trace as State puts the collected trace before the long-context block on a fresh pass instead of appending it after. On GraphWalks Parents, DeepSeek V4 Pro Preview goes from 29.2% on the initial pass and 43.0% with the matched append control to 81.8%, and GLM-5.2 goes from 66.4% and 83.2% to 100.0%. It wins in 26 of 27 reported combinations of model, task, and metric with no architecture change.

65Memory
CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?

CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?

Damien Sileo and Dimitri Kachler introduce CordisBench, a 1,200-question benchmark for a reasoning burden that only appears once agents can rewrite the software running them: predicting how a local component change propagates through dependencies and teardown.

66Agents
Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

Milad Rezaei Hajidehi, Qitong Wang and Stratos Idreos (Harvard) propose agentic data cracking, where a sub-agent forks from an already-loaded document context to speculatively extract structure that future queries will reuse.

67Agents
Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search

Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search

Yuan Chang and Xiaoqi Chen show that a single-lineage prompt optimizer with rollout feedback matches or beats GEPA using fewer rollouts, and that the gap widens as the teacher model gets stronger.

68Reasoning
Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture

Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture

Francisco Arrabal-Campos and colleagues (University of Almeria) build a minimal complete cognitive architecture with a recurrent reasoner, adaptive halting, and a value module, then ask of each part whether the function emerges from gradient descent or has to be computed explicitly.

69Reasoning
Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for AI Co-Scientists in Protein Characterization Workflows

Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for AI Co-Scientists in Protein Characterization Workflows

Maia Kapur and colleagues run a controlled ablation on a production agentic science platform, using protein function characterization as a verifiable task to separate what federation topology, harness type, model choice, and prompt expertise each contribute.

70Reasoning
ReasoningMathematical discoveries from program search with large language models

Mathematical discoveries from program search with large language models

Evolutionary search over programs an LLM writes, filtered by a systematic evaluator, produced new results in extremal combinatorics. It is the direct predecessor of AlphaEvolve and the first case of a language model making a discovery on an established open problem.

71Reasoning
Large Language Models Cannot Self-Correct Reasoning Yet

Large Language Models Cannot Self-Correct Reasoning Yet

Without external feedback, a model asked to correct its own reasoning does not improve and often gets worse. This is the empirical limit that separates a self-improvement loop with a real signal from one talking to itself.

72Reasoning
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026