AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit Signal
Yunxiang Mo, Donghao Zhao (HKUST) and Hejia Geng (University of Oxford) preregister a sweep of 3,520 self-consensus early-exit rules and find that none clears three acceptance gates, because agreement measures answer persistence rather than reasoning termination.

TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards
Rui Sun, Zhan Shi and Bing He (independent researchers) train diagnostic reasoning agents by sampling an intervention, injecting it into a simulator, and generating the observations it would produce, so the hidden intervention supplies an oracle label for a task where real ground truth would require expert investigation.

Prompt Repetition Improves Non-Reasoning LLMs
Yaniv Leviathan, Matan Kalman and Yossi Matias at Google Research report that simply repeating the input prompt improves non-reasoning model performance across Gemini, GPT, Claude and DeepSeek.

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
Seogyeong Jeong and colleagues at KAIST and NAVER AI Lab test whether the functional operations inside a chain of thought, such as problem formulation, goal decomposition and deduction, have distinct geometric structure in hidden representations.

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Ji Soo Lee and colleagues at Meta and KAIST build WearableQA from the wearable time series, blood biomarkers, and demographics of 200 real users with up to 500 days of daily measurements each.

LOCI: A Locator-Critic with Refinement Loop
Walid Bousselham, Mathilde Caron, Arsha Nagrani and Cordelia Schmid at Google DeepMind argue that VLM failures on hard visual tasks come from failing to locate the relevant detail, not from weak high-level reasoning, and fix it with a two-agent loop that needs no training.

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
Jacqueline He and colleagues at Meta AI, the University of Washington and Princeton show that standard knowledge distillation helps reasoning and hurts factual recall during mid-training, and trace the cause to teacher confidence.

What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation
Daisuke Kikuta (NTT) studies revision propagation, where a user asks for one local change and the model must find and update every dependent part of an artifact whose dependencies are buried in the conversation history.

LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL
Sijie Wang, Zhiqiang Tan, Xinrui Yang and Shaohuai Shi at Harbin Institute of Technology Shenzhen remove the recomputation step that DanceGRPO and FlowGRPO perform after rollout, which is mathematically redundant when rollout and update share a backend under on-policy training.

RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory
Yuxiang Wang and colleagues fix two limitations of looped-layer latent recurrence at once, letting each iteration attend to its own earlier states and letting the model decide how many loops a given input deserves.

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories
Yigit Utku Bulut supplies the two counterfactual controls that the breakthrough-moment and early-legible-fate readings of reasoning traces have been missing, and both readings largely fail to survive them.

GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving
Qiankun Ma and colleagues point out that every KV compression method fixes the per-request budget in advance and only decides what to keep, then make capacity itself a runtime resource that grows on demand.

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards
Leqi Zheng and colleagues propose Gradient-Aligned Reward, which builds a dense reasoning-aware reward by comparing each rollout's gradient direction to an expert-anchor gradient, using expert solutions already sitting in the training corpus.

Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning
Kevin Du, Alexander Hoyle, Laura Ruis and Acyr Locatelli (ETH Zurich, work done during an internship at Cohere) test whether the text of a reasoning step actually encodes how much that step mattered, using Monte Carlo advantage as ground truth, and find only partial recoverability.

RuleMem: Active Rule Memory for Long-Term Conversational Agents
Xingyuan Zeng and colleagues propose RuleMem, which induces reusable natural-language Horn clauses from conversation history so that agent memory actively guides retrieval and reasoning instead of sitting as passively stored facts.

Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
Heng Wang and colleagues at Salesforce AI Research, UIUC and Cornell show that the token-importance scores every KV cache evictor computes are close to worthless: evicting uniformly at random inside each head matches the strongest prior evictor while serving 32 to 43% higher throughput.

Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers
Where you put a reasoning trace changes long-context accuracy by up to 50 points. Transformers process causally, so a task state discovered late cannot guide reading that already happened, and Trace as State puts the collected trace before the long-context block on a fresh pass instead of appending it after. On GraphWalks Parents, DeepSeek V4 Pro Preview goes from 29.2% on the initial pass and 43.0% with the matched append control to 81.8%, and GLM-5.2 goes from 66.4% and 83.2% to 100.0%. It wins in 26 of 27 reported combinations of model, task, and metric with no architecture change.

CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?
Damien Sileo and Dimitri Kachler introduce CordisBench, a 1,200-question benchmark for a reasoning burden that only appears once agents can rewrite the software running them: predicting how a local component change propagates through dependencies and teardown.

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
Milad Rezaei Hajidehi, Qitong Wang and Stratos Idreos (Harvard) propose agentic data cracking, where a sub-agent forks from an already-loaded document context to speculatively extract structure that future queries will reuse.

Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search
Yuan Chang and Xiaoqi Chen show that a single-lineage prompt optimizer with rollout feedback matches or beats GEPA using fewer rollouts, and that the gap widens as the teacher model gets stronger.

Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture
Francisco Arrabal-Campos and colleagues (University of Almeria) build a minimal complete cognitive architecture with a recurrent reasoner, adaptive halting, and a value module, then ask of each part whether the function emerges from gradient descent or has to be computed explicitly.

Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for AI Co-Scientists in Protein Characterization Workflows
Maia Kapur and colleagues run a controlled ablation on a production agentic science platform, using protein function characterization as a verifiable task to separate what federation topology, harness type, model choice, and prompt expertise each contribute.
Mathematical discoveries from program search with large language models
Evolutionary search over programs an LLM writes, filtered by a systematic evaluator, produced new results in extremal combinatorics. It is the direct predecessor of AlphaEvolve and the first case of a language model making a discovery on an established open problem.

Large Language Models Cannot Self-Correct Reasoning Yet
Without external feedback, a model asked to correct its own reasoning does not improve and often gets worse. This is the empirical limit that separates a self-improvement loop with a real signal from one talking to itself.