🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
420 papers · ReasoningClear filters →
Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

Xing Zhang, Peiyang He and colleagues at AWS Forward Deployed Engineering make the verifier itself the evolving object in a self-improving agent loop, building inspectable graders from small deterministic drawback detectors. Accepted at the NeurIPS 2026 workshop Who Verifies the Agents?

01Agents
Thinking Inertia: LLMs Keep Thinking When Told Not To

Thinking Inertia: LLMs Keep Thinking When Told Not To

Dianqiao Lei (Tsinghua), Kevin Qinghong Lin and Philip Torr (Oxford), and Pan Lu and James Zou (Stanford) show that LLMs keep producing explicit reasoning when instructed not to, a behavior they call Thinking Inertia. Accepted at NeurIPS 2026.

02Reasoning
Finding Blind Spots in AppWorld and WorkArena Task Verifiers

Finding Blind Spots in AppWorld and WorkArena Task Verifiers

Richard Abrich (OpenAdapt.AI) audits the shipped task verifiers of AppWorld and WorkArena for false accepts using source-informed mutation tests. Accepted at a NeurIPS 2026 workshop on verifying agents.

03Reasoning
Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models

Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models

Xingang Guo, Jing Gu, Jared Lichtarge and colleagues at Scale AI (with Elorian) introduce Humanity's Sixth Sense (HSS), a benchmark for the intuitive visual reasoning people perform at a glance.

04Multimodal
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

Zewei Zhou, Boris Ivanovic, Marco Pavone and colleagues at NVIDIA with UCLA, UC Berkeley and Stanford introduce VeriFine, a harness that improves the policy, its training curriculum and its judge together for embodied reasoning.

05Robotics
SanSi: A Looped Typed Decision Model for System 1.5 Thinking

SanSi: A Looped Typed Decision Model for System 1.5 Thinking

Shuyu Gan, Young-Jun Lee and Dongyeop Kang at the University of Minnesota present SanSi, a typed decision model that loops the same layers several times before a single option readout, a regime between one-pass decisions and generated reasoning that they call System 1.5 thinking.

06Reasoning
CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

Yifan Zhang, Yutong Dai, Ran Xu, Zeyuan Chen and colleagues at Salesforce AI Research introduce CLIFT, which trains a web agent to verify its own rollouts and reuses that verifier for test-time trajectory selection without an external judge.

07Agents
Continual Graph Memory for Mathematical Research Agents

Continual Graph Memory for Mathematical Research Agents

Junyi Zhang, Jinxi Yu, Eric Hanchen Jiang and colleagues at UCLA, with senior authors including Kai-Wei Chang, Raghu Meka, Nanyun Peng, Amit Sahai, Terence Tao and Wei Wang, present Ansatz, a mathematical research agent built around Continual Graph Memory, which stores proof progress as typed graphs instead of flat text.

08Memory
Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents

Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents

Ankur Samanta, Kaveh Hassani and Anirudh Goyal at Meta AI, with Yonathan Efroni (Tel Aviv) and Paul Sajda (Columbia), introduce MIRA, a research-agent architecture in which an outer meta-reasoner decides what to investigate next and a fresh executor carries out each investigation, and train that outer policy with RL.

09Reasoning
RLTL;DR: Self-improvement by Internalizing Self-generated Feedback

RLTL;DR: Self-improvement by Internalizing Self-generated Feedback

Michael Kirchhof, Andrew Szot, Alexander Toshev and colleagues at Apple introduce RLTL;DR, an RL recipe for tasks where the policy never succeeds: after each failed attempt the policy writes a one-line insight from the verifier output, later rollouts are conditioned on those insights, and the insight tokens are trained on so the lesson moves into the weights.

10Reasoning
Rational Clarification by Assistive Agents via Value-of-Information Reasoning

Rational Clarification by Assistive Agents via Value-of-Information Reasoning

T. Duy Nguyen-Hien and Wee Sun Lee (NUS), Yee Whye Teh (Oxford) and Tan Zhi-Xuan introduce REVOIR, an inference-time method that decides whether an assistant should ask a clarifying question by estimating how much the answer would raise expected task reward, net of the cost of asking.

11Reasoning
VISTA: A Visual Harness for Reasoning in an Interactive World

VISTA: A Visual Harness for Reasoning in an Interactive World

Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu and Kaiming He at MIT introduce VISTA, a visual harness that lets a general-purpose multimodal model act in interactive environments from image observations, with a lossless visual memory it can search and rearrange while it reasons.

12Reasoning
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Minki Kang (KAIST, intern at NVIDIA), Byung-Kwan Lee, Yu-Chiang Frank Wang and colleagues at NVIDIA introduce Mid-Harness, which spends test-time compute at the boundary between model and harness: it samples several candidate shell actions and verifies them before one is executed.

13Agents
Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

Paras Dahal, Anton Bakhtin, Taco Cohen, Rob Fergus, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Anirudh Goyal and colleagues at Meta Superintelligence Labs introduce agentic meta-reasoning, an inference-time harness in which a controller decides what work to assign, so that choosing the next step becomes a structured reasoning process separate from the task work.

14Reasoning
SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

Jianxing Chen, Xiao Yu, Shipra Agrawal and Zhou Yu at Columbia University introduce SCOUT, a two-stage safety verifier for computer-use agents that writes task-specific rubrics and then probes the post-execution environment for evidence of harm.

15Agents
HeurEvo: Agentic Evolution of Hybrid Solver-Augmented Heuristics for Time-Critical Mathematical Optimization

HeurEvo: Agentic Evolution of Hybrid Solver-Augmented Heuristics for Time-Critical Mathematical Optimization

Feijie Wu, Hugo Barbalho, Konstantina Mellou and colleagues at Microsoft Research and Purdue introduce HeurEvo, which co-evolves the plan, code and reusable components of hybrid heuristics that call mathematical-programming solvers under tight runtime limits.

16Agents
SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

Xinyue Zeng and colleagues at Virginia Tech, UW-Madison and Dartmouth propose SAGE, which adds structural guidance to long-horizon reasoning to counter exploration and compounding biases under sparse rewards.

17Reasoning
When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression

When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression

Mingxuan Wang and colleagues (TierFlow team, Renmin University Gaoling School) study when an agent can safely drop its earlier reasoning, and propose ICLR (Interaction Aware Compression for Long Horizon Reasoning), a training-free online method.

18Reasoning
SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

SkillGym turns human-written agent skills into 2,756 verifiable training environments across 12 categories, each with code-based checkers, and collects 8,364 successful trajectories for fine-tuning. Under Claude Code, fine-tuning Qwen3.5-35B-A3B adds 19.10 points on Terminal-Bench 2.1 and 28.13 points on skill-assisted SkillsBench v1.1, where it reaches 51.47%, above the reported scores for Claude Sonnet 4.6 and GPT-5.4 Mini. With no skills loaded, the trained model still beats the base model that has the skills in context.

19Agents
Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning

Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning

Ismail Labiad, Rémi Munos, Julia Kempe and colleagues at Meta FAIR with Université Paris-Saclay and NYU train a small concept generator with RL so that the hints it writes raise the pass@k of a larger frozen answer model.

20Reasoning
Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Many agentic memory systems use an autoregressive LLM to decide how memories are organized, retrieved, and used, which puts expensive generation on the critical path of every memory operation. Jev-Mem borrows its design from System-One/System-Two cognition and hands those decisions to a lightweight controller.

21Agents
Self-Organizing Agent Teams Learn to Reason Together

Self-Organizing Agent Teams Learn to Reason Together

Multi-agent systems usually fix roles and protocols in advance. Researchers from Stanford and Together AI let a fixed team of models learn how to organize its own collaboration from past exchanges.

22Agents
WFM: Wiki Foundation Model for Complex Agentic Reasoning

WFM: Wiki Foundation Model for Complex Agentic Reasoning

More agents now store long-term memory as an LLM Wiki, a folder of markdown pages linked to each other. Each page holds dense text and the links hold structure, and WFM is a Wiki Foundation Model trained to use both when retrieving.

23Agents
Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition

Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition

Dohun Lee and Hyunwoo Park (Seoul National University) measure structural and intent faithfulness of LLM pricing agents in Bertrand competition and find both are unrelated to whether the agents collude.

24Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026