AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
WFM: Wiki Foundation Model for Complex Agentic Reasoning
More agents now store long-term memory as an LLM Wiki, a folder of markdown pages linked to each other. Each page holds dense text and the links hold structure, and WFM is a Wiki Foundation Model trained to use both when retrieving.

Unlocking the Unsolvable: Teacher-Guided Curriculum for Data-Efficient RLVR
Zhu (independent) and Han (Amazon, work done independently) show that RLVR problems a model cannot yet solve, which normally yield zero learning signal, can be made useful by giving partial teacher reasoning traces and withdrawing them step by step.

Efficiently Linking Unstructured Data for Multi-step Reasoning
Jiaming Liang, Haydn Jones, Jacob R. Gardner, Mark Yatskar and Zachary Ives (University of Pennsylvania) build a query engine for the retrieval step that sits under agentic reasoning pipelines, executing filters, multi-vector search, relational joins and similarity joins together.

Language-model groups overstate consensus when replaying human deliberation on a reasoning task
Tengfei Shao replays 100 held-out human Wason group discussions with matched LLM agent groups and finds the agent groups reach full consensus far more often than the people they stand in for.

LLM-as-an-Improver: Turning Verification into Better Candidates
Akiyoshi Tomihari and Yuma Ichikawa ask whether verifier feedback can improve the candidate pool rather than only rank it, and propose Verify-Repair-Reselect.

AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation
Keshu Wu and colleagues at Texas A&M and collaborators treat air-ground co-simulation scenario generation as compilation with verification, so a scenario that runs is also checked against the relationships the user asked for.

CovR: Coverage-Aware Hardware Verification via Reasoning-Guided Reinforcement Learning
Abdelatty, Nouh and Reda (Brown University) build CovR, an agentic testbench-generation system for RTL hardware verification that optimizes for coverage rather than functional correctness alone, and distill the resulting behavior into a student model with simulation-derived rewards.

Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes
Qirui Chen, Renjie Pi, Jiahui Gao and Lingpeng Kong (Zhejiang University, the University of Hong Kong and HKUST) convert failed reasoning rollouts into recovery training data, so a model learns to continue correctly from an already-wrong intermediate state.

PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces
Pyrros Koussios and colleagues introduce PetriBench, which evaluates LLM reasoning over dynamic state spaces using Petri nets, a formalism for concurrent and distributed systems, with exact ground truth and procedural generation.

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning
Wenjie Liao, Liangjie Zhao and Zehong Cao train task generation, execution and evaluation jointly in UnifiedPlayers, rather than pairing self-generated trajectories with a static verifier.

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
Jaejun Shim and colleagues treat efficient reasoning as instance-adaptive compute allocation and train When2Think to choose between direct answering and extended reasoning per problem.

Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
Caiqi Zhang, Nigel Collier, Dharshan Kumaran and colleagues at the University of Cambridge and Google DeepMind propose XConf, which estimates an LLM's confidence from a record of its own graded past episodes rather than from the current inference alone.

Metacognitive Steering: Learning the Structure of Scientific Judgment
Vincent Karpf and colleagues (Autopoiesis Sciences) find a low-dimensional control structure in Kimi 2.6 that corresponds to scientific judgment and use it to steer reasoning at inference time.

Available but Unclaimed: An Empirical Study of Human-AI Synergy
Robin Welsch, Albrecht Schmidt and colleagues (Aalto, LMU) ran a 535-person study of reasoning with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash or Kimi K3, and measured how much model accuracy reaches the human-AI team.

Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA
Kailong Fan, Yichen Wu and colleagues (Harvard Medical School/MGH) show why majority-vote test-time RL collapses on medical multiple-choice QA and propose PROSE, which rewards reasoning steps instead of answer agreement.

Verifiable Social Reasoning for LLM Assistants
People ask assistants for social advice constantly, and the assistant only hears the user's version of events, which makes it hard to check whether it read the situation correctly. Google Research builds that ground truth by simulation, with a target agent holding a hidden motive while a user agent relays events to the assistant, which then has to infer the motive. Across 24k human annotations validating the simulations and 12 LLMs tested, biased framing from the user shifted the assistant's answer, and longer conversations with room for clarifying questions did not reliably help.

Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering
Alexander Krentsel, Shubham Agarwal, Mert Cemri, Shu Liu and colleagues at UC Berkeley, including Matei Zaharia and Ion Stoica, argue that passing tests or even a proof cannot guarantee acceptable deployed behavior, and propose an outer loop that revises requirements, environment models and evaluators from deployment evidence.

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
Utkarsh Soni and colleagues at Manulife build TAM, a benchmark of real tasks that require following manuals with tens of thousands of rules, in ICD-10-CM clinical coding and U.S. federal sentencing.

MindTopo: Can Foundation Models Reason in Topological Space?
Yunfei Ge, Manling Li and colleagues (Northwestern University, Microsoft Research and Stanford) introduce MindTopo, a benchmark of topological reasoning and planning, and find every tested multimodal model does worse when it has to act on topological relations than when it only identifies them.

Beyond Solver Verdicts: Generative Reward Models for Autoformalization
Vikash Singh, Debargha Ganguly and colleagues (Case Western Reserve University with Amazon Web Services) show that solver verdicts cannot detect formal translations that are wrong but still return the expected verdict, and train a generative verifier that scores equivalence to the reference formalization.

Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation
Dong Li, Biqing Qi and colleagues (Shanghai AI Laboratory with Harbin Institute of Technology and others) build ARCHE, an agent that proposes reaction mechanisms, runs computational chemistry workflows to test them and revises conclusions from the computed evidence.

When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making
Ken Chen, Saman Halgamuge and colleagues (University of Melbourne) resolve disagreement between LLM agents by comparing each agent's forward answer with a posterior computed by Bayesian backward reasoning, which is less likely to share the same errors.

Negative Self-Distillation: Learning to Reason by Avoiding Flaws
Rongcan Pei, Yu Meng and colleagues (University of Virginia) replace on-policy self-distillation's imitation of privileged solutions with Negative Self-Distillation, which pushes the model away from a self-generated flawed reasoner and needs no ground-truth answers.

Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification
Joshua Ong Jun Leang and colleagues (MBZUAI, Imperial, Edinburgh and UCL) build Magenta, a training-free agent that answers a math problem, restates the answer in Lean 4 and produces a machine-checked proof, reaching 100% on recent olympiad benchmarks.