AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
Xing Zhang, Peiyang He and colleagues at AWS Forward Deployed Engineering make the verifier itself the evolving object in a self-improving agent loop, building inspectable graders from small deterministic drawback detectors. Accepted at the NeurIPS 2026 workshop Who Verifies the Agents?

Thinking Inertia: LLMs Keep Thinking When Told Not To
Dianqiao Lei (Tsinghua), Kevin Qinghong Lin and Philip Torr (Oxford), and Pan Lu and James Zou (Stanford) show that LLMs keep producing explicit reasoning when instructed not to, a behavior they call Thinking Inertia. Accepted at NeurIPS 2026.

Finding Blind Spots in AppWorld and WorkArena Task Verifiers
Richard Abrich (OpenAdapt.AI) audits the shipped task verifiers of AppWorld and WorkArena for false accepts using source-informed mutation tests. Accepted at a NeurIPS 2026 workshop on verifying agents.

Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models
Xingang Guo, Jing Gu, Jared Lichtarge and colleagues at Scale AI (with Elorian) introduce Humanity's Sixth Sense (HSS), a benchmark for the intuitive visual reasoning people perform at a glance.

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
Zewei Zhou, Boris Ivanovic, Marco Pavone and colleagues at NVIDIA with UCLA, UC Berkeley and Stanford introduce VeriFine, a harness that improves the policy, its training curriculum and its judge together for embodied reasoning.

SanSi: A Looped Typed Decision Model for System 1.5 Thinking
Shuyu Gan, Young-Jun Lee and Dongyeop Kang at the University of Minnesota present SanSi, a typed decision model that loops the same layers several times before a single option readout, a regime between one-pass decisions and generated reasoning that they call System 1.5 thinking.

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling
Yifan Zhang, Yutong Dai, Ran Xu, Zeyuan Chen and colleagues at Salesforce AI Research introduce CLIFT, which trains a web agent to verify its own rollouts and reuses that verifier for test-time trajectory selection without an external judge.

Continual Graph Memory for Mathematical Research Agents
Junyi Zhang, Jinxi Yu, Eric Hanchen Jiang and colleagues at UCLA, with senior authors including Kai-Wei Chang, Raghu Meka, Nanyun Peng, Amit Sahai, Terence Tao and Wei Wang, present Ansatz, a mathematical research agent built around Continual Graph Memory, which stores proof progress as typed graphs instead of flat text.

Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents
Ankur Samanta, Kaveh Hassani and Anirudh Goyal at Meta AI, with Yonathan Efroni (Tel Aviv) and Paul Sajda (Columbia), introduce MIRA, a research-agent architecture in which an outer meta-reasoner decides what to investigate next and a fresh executor carries out each investigation, and train that outer policy with RL.

RLTL;DR: Self-improvement by Internalizing Self-generated Feedback
Michael Kirchhof, Andrew Szot, Alexander Toshev and colleagues at Apple introduce RLTL;DR, an RL recipe for tasks where the policy never succeeds: after each failed attempt the policy writes a one-line insight from the verifier output, later rollouts are conditioned on those insights, and the insight tokens are trained on so the lesson moves into the weights.

Rational Clarification by Assistive Agents via Value-of-Information Reasoning
T. Duy Nguyen-Hien and Wee Sun Lee (NUS), Yee Whye Teh (Oxford) and Tan Zhi-Xuan introduce REVOIR, an inference-time method that decides whether an assistant should ask a clarifying question by estimating how much the answer would raise expected task reward, net of the cost of asking.

VISTA: A Visual Harness for Reasoning in an Interactive World
Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu and Kaiming He at MIT introduce VISTA, a visual harness that lets a general-purpose multimodal model act in interactive environments from image observations, with a lossless visual memory it can search and rearrange while it reasons.

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Minki Kang (KAIST, intern at NVIDIA), Byung-Kwan Lee, Yu-Chiang Frank Wang and colleagues at NVIDIA introduce Mid-Harness, which spends test-time compute at the boundary between model and harness: it samples several candidate shell actions and verifies them before one is executed.

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
Paras Dahal, Anton Bakhtin, Taco Cohen, Rob Fergus, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Anirudh Goyal and colleagues at Meta Superintelligence Labs introduce agentic meta-reasoning, an inference-time harness in which a controller decides what work to assign, so that choosing the next step becomes a structured reasoning process separate from the task work.

SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety
Jianxing Chen, Xiao Yu, Shipra Agrawal and Zhou Yu at Columbia University introduce SCOUT, a two-stage safety verifier for computer-use agents that writes task-specific rubrics and then probes the post-execution environment for evidence of harm.

HeurEvo: Agentic Evolution of Hybrid Solver-Augmented Heuristics for Time-Critical Mathematical Optimization
Feijie Wu, Hugo Barbalho, Konstantina Mellou and colleagues at Microsoft Research and Purdue introduce HeurEvo, which co-evolves the plan, code and reusable components of hybrid heuristics that call mathematical-programming solvers under tight runtime limits.

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance
Xinyue Zeng and colleagues at Virginia Tech, UW-Madison and Dartmouth propose SAGE, which adds structural guidance to long-horizon reasoning to counter exploration and compounding biases under sparse rewards.

When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression
Mingxuan Wang and colleagues (TierFlow team, Renmin University Gaoling School) study when an agent can safely drop its earlier reasoning, and propose ICLR (Interaction Aware Compression for Long Horizon Reasoning), a training-free online method.

SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving
SkillGym turns human-written agent skills into 2,756 verifiable training environments across 12 categories, each with code-based checkers, and collects 8,364 successful trajectories for fine-tuning. Under Claude Code, fine-tuning Qwen3.5-35B-A3B adds 19.10 points on Terminal-Bench 2.1 and 28.13 points on skill-assisted SkillsBench v1.1, where it reaches 51.47%, above the reported scores for Claude Sonnet 4.6 and GPT-5.4 Mini. With no skills loaded, the trained model still beats the base model that has the skills in context.

Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning
Ismail Labiad, Rémi Munos, Julia Kempe and colleagues at Meta FAIR with Université Paris-Saclay and NYU train a small concept generator with RL so that the hints it writes raise the pass@k of a larger frozen answer model.

Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents
Many agentic memory systems use an autoregressive LLM to decide how memories are organized, retrieved, and used, which puts expensive generation on the critical path of every memory operation. Jev-Mem borrows its design from System-One/System-Two cognition and hands those decisions to a lightweight controller.

Self-Organizing Agent Teams Learn to Reason Together
Multi-agent systems usually fix roles and protocols in advance. Researchers from Stanford and Together AI let a fixed team of models learn how to organize its own collaboration from past exchanges.

WFM: Wiki Foundation Model for Complex Agentic Reasoning
More agents now store long-term memory as an LLM Wiki, a folder of markdown pages linked to each other. Each page holds dense text and the links hold structure, and WFM is a Wiki Foundation Model trained to use both when retrieving.

Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition
Dohun Lee and Hyunwoo Park (Seoul National University) measure structural and intent faithfulness of LLM pricing agents in Bertrand competition and find both are unrelated to whether the agents collude.