🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
420 papers · ReasoningClear filters →
WFM: Wiki Foundation Model for Complex Agentic Reasoning

WFM: Wiki Foundation Model for Complex Agentic Reasoning

More agents now store long-term memory as an LLM Wiki, a folder of markdown pages linked to each other. Each page holds dense text and the links hold structure, and WFM is a Wiki Foundation Model trained to use both when retrieving.

25Agents
Unlocking the Unsolvable: Teacher-Guided Curriculum for Data-Efficient RLVR

Unlocking the Unsolvable: Teacher-Guided Curriculum for Data-Efficient RLVR

Zhu (independent) and Han (Amazon, work done independently) show that RLVR problems a model cannot yet solve, which normally yield zero learning signal, can be made useful by giving partial teacher reasoning traces and withdrawing them step by step.

26Reasoning
Efficiently Linking Unstructured Data for Multi-step Reasoning

Efficiently Linking Unstructured Data for Multi-step Reasoning

Jiaming Liang, Haydn Jones, Jacob R. Gardner, Mark Yatskar and Zachary Ives (University of Pennsylvania) build a query engine for the retrieval step that sits under agentic reasoning pipelines, executing filters, multi-vector search, relational joins and similarity joins together.

27Reasoning
Language-model groups overstate consensus when replaying human deliberation on a reasoning task

Language-model groups overstate consensus when replaying human deliberation on a reasoning task

Tengfei Shao replays 100 held-out human Wason group discussions with matched LLM agent groups and finds the agent groups reach full consensus far more often than the people they stand in for.

28Agents
LLM-as-an-Improver: Turning Verification into Better Candidates

LLM-as-an-Improver: Turning Verification into Better Candidates

Akiyoshi Tomihari and Yuma Ichikawa ask whether verifier feedback can improve the candidate pool rather than only rank it, and propose Verify-Repair-Reselect.

29Reasoning
AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation

AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation

Keshu Wu and colleagues at Texas A&M and collaborators treat air-ground co-simulation scenario generation as compilation with verification, so a scenario that runs is also checked against the relationships the user asked for.

30Agents
CovR: Coverage-Aware Hardware Verification via Reasoning-Guided Reinforcement Learning

CovR: Coverage-Aware Hardware Verification via Reasoning-Guided Reinforcement Learning

Abdelatty, Nouh and Reda (Brown University) build CovR, an agentic testbench-generation system for RTL hardware verification that optimizes for coverage rather than functional correctness alone, and distill the resulting behavior into a student model with simulation-derived rewards.

31Reasoning
Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes

Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes

Qirui Chen, Renjie Pi, Jiahui Gao and Lingpeng Kong (Zhejiang University, the University of Hong Kong and HKUST) convert failed reasoning rollouts into recovery training data, so a model learns to continue correctly from an already-wrong intermediate state.

32Reasoning
PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces

PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces

Pyrros Koussios and colleagues introduce PetriBench, which evaluates LLM reasoning over dynamic state spaces using Petri nets, a formalism for concurrent and distributed systems, with exact ground truth and procedural generation.

33Reasoning
UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

Wenjie Liao, Liangjie Zhao and Zehong Cao train task generation, execution and evaluation jointly in UnifiedPlayers, rather than pairing self-generated trajectories with a static verifier.

34Reasoning
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

Jaejun Shim and colleagues treat efficient reasoning as instance-adaptive compute allocation and train When2Think to choose between direct answering and extended reasoning per problem.

35Reasoning
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

Caiqi Zhang, Nigel Collier, Dharshan Kumaran and colleagues at the University of Cambridge and Google DeepMind propose XConf, which estimates an LLM's confidence from a record of its own graded past episodes rather than from the current inference alone.

36Agents
Metacognitive Steering: Learning the Structure of Scientific Judgment

Metacognitive Steering: Learning the Structure of Scientific Judgment

Vincent Karpf and colleagues (Autopoiesis Sciences) find a low-dimensional control structure in Kimi 2.6 that corresponds to scientific judgment and use it to steer reasoning at inference time.

37Reasoning
Available but Unclaimed: An Empirical Study of Human-AI Synergy

Available but Unclaimed: An Empirical Study of Human-AI Synergy

Robin Welsch, Albrecht Schmidt and colleagues (Aalto, LMU) ran a 535-person study of reasoning with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash or Kimi K3, and measured how much model accuracy reaches the human-AI team.

38Reasoning
Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA

Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA

Kailong Fan, Yichen Wu and colleagues (Harvard Medical School/MGH) show why majority-vote test-time RL collapses on medical multiple-choice QA and propose PROSE, which rewards reasoning steps instead of answer agreement.

39Reinforcement Learning
Verifiable Social Reasoning for LLM Assistants

Verifiable Social Reasoning for LLM Assistants

People ask assistants for social advice constantly, and the assistant only hears the user's version of events, which makes it hard to check whether it read the situation correctly. Google Research builds that ground truth by simulation, with a target agent holding a hidden motive while a user agent relays events to the assistant, which then has to infer the motive. Across 24k human annotations validating the simulations and 12 LLMs tested, biased framing from the user shifted the assistant's answer, and longer conversations with room for clarifying questions did not reliably help.

40Evaluation
Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering

Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering

Alexander Krentsel, Shubham Agarwal, Mert Cemri, Shu Liu and colleagues at UC Berkeley, including Matei Zaharia and Ion Stoica, argue that passing tests or even a proof cannot guarantee acceptable deployed behavior, and propose an outer loop that revises requirements, environment models and evaluators from deployment evidence.

41Agents
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

Utkarsh Soni and colleagues at Manulife build TAM, a benchmark of real tasks that require following manuals with tens of thousands of rules, in ICD-10-CM clinical coding and U.S. federal sentencing.

42Reasoning
MindTopo: Can Foundation Models Reason in Topological Space?

MindTopo: Can Foundation Models Reason in Topological Space?

Yunfei Ge, Manling Li and colleagues (Northwestern University, Microsoft Research and Stanford) introduce MindTopo, a benchmark of topological reasoning and planning, and find every tested multimodal model does worse when it has to act on topological relations than when it only identifies them.

43Reasoning
Beyond Solver Verdicts: Generative Reward Models for Autoformalization

Beyond Solver Verdicts: Generative Reward Models for Autoformalization

Vikash Singh, Debargha Ganguly and colleagues (Case Western Reserve University with Amazon Web Services) show that solver verdicts cannot detect formal translations that are wrong but still return the expected verdict, and train a generative verifier that scores equivalence to the reference formalization.

44Reinforcement Learning
Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation

Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation

Dong Li, Biqing Qi and colleagues (Shanghai AI Laboratory with Harbin Institute of Technology and others) build ARCHE, an agent that proposes reaction mechanisms, runs computational chemistry workflows to test them and revises conclusions from the computed evidence.

45Agents
When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making

When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making

Ken Chen, Saman Halgamuge and colleagues (University of Melbourne) resolve disagreement between LLM agents by comparing each agent's forward answer with a posterior computed by Bayesian backward reasoning, which is less likely to share the same errors.

46Agents
Negative Self-Distillation: Learning to Reason by Avoiding Flaws

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

Rongcan Pei, Yu Meng and colleagues (University of Virginia) replace on-policy self-distillation's imitation of privileged solutions with Negative Self-Distillation, which pushes the model away from a self-generated flawed reasoner and needs no ground-truth answers.

47Training
Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification

Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification

Joshua Ong Jun Leang and colleagues (MBZUAI, Imperial, Edinburgh and UCL) build Magenta, a training-free agent that answers a math problem, restates the answer in Lean 4 and produces a machine-checked proof, reaching 100% on recent olympiad benchmarks.

48Reasoning
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026