AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
LOCI: A Locator-Critic with Refinement Loop
Walid Bousselham, Mathilde Caron, Arsha Nagrani and Cordelia Schmid at Google DeepMind argue that VLM failures on hard visual tasks come from failing to locate the relevant detail, not from weak high-level reasoning, and fix it with a two-agent loop that needs no training.

SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation
Qi Liu, Qinzheng Wang and Yiming Bie build SimSkill, a self-evolving agent over the SUMO traffic simulator that finds its own capability gaps, writes and solves grounded tasks, and consolidates the results into episodic, procedural and semantic memory without touching the backbone weights.

Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection
Weijie Liu and colleagues at HKU build Dude, a dual-detection multi-agent system for finding places where a paper's claims and its released code disagree, and diagnose why naive multi-agent designs over-report.

Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation
Yan Tang and colleagues formalize proactive service as a partially observable sequential decision process constrained by authorization and risk, where staying silent is a first-class action with option value.

Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation
Xuanfa Jin and colleagues at CASIA and UCL attack the shared-misconception failure in multi-agent debate with R2-MAD, giving debating agents an experience memory from past debates plus per-agent confidence weights.

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models
Xingming Long and colleagues introduce NTEP, an annotation scheme that names the necessary external evidence and the tool calls that must produce it, and NTEP-R, a reward that pays the agent per tool call rather than only on the final answer.

Bioinfoysis Technical Report
The DeepAutonomy Team introduces Bioinfoysis, a multi-agent harness that treats a bioinformatics request as a persistent analysis run whose conclusions stay attached to the artifacts that produced them, reaching 82.4% on BixBench.

The Natural Language Interaction Protocol and Standard for AI Agents
Luyi Xing and a cross-industry group present NLIP, an application-layer protocol for AI-agent interaction standardized by Ecma International, aimed at the interoperability gap that MCP and A2A only partly cover.

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
Yaxing Lyu and colleagues build KC-Bench to measure whether a tool-using model can reconcile user instructions, its own parametric knowledge and live environmental observations before it acts on any of them.

Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning
A large InstaDeep team with AIMS and Stellenbosch extends offline sequence models to variable agent counts and multi-task observation and action spaces, then measures which scaling axis actually produces zero-shot transfer in offline multi-agent RL.

Efficient Test-Time Adaptation through Human-AI Interaction
Zora Zhiruo Wang, Apurva Gandhi, Rulin Shao and colleagues at Carnegie Mellon University, the University of Washington and Handshake propose TAHI, which turns the interaction history between one professional and their agent into both context and weight updates, plus an evolving per-user rubric that encodes the criteria the user never wrote down.

TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents
Jinwei Gan at Nanjing University introduces TIGPO, which keeps a persistent per-task transition graph across policy updates so that credit assignment for long-horizon agents can draw on transitions discovered by earlier policy versions rather than only the current batch.

Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
Zhaoyuan Huang and colleagues at Shanghai Jiao Tong and Ant Group ask whether GUI agents know when not to act, build CONFLICTGUI to measure it, and find severe execution-biased overcompliance across five widely used agents.

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
Jie Wu and colleagues on the Qwen team at Alibaba with Tsinghua turn the pile of existing terminal-agent trajectories into executable environments, on the observation that a trajectory's tool-execution history already exposes the environment it ran in.

A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms
Google DeepMind ran a research collective of 100 autonomous LLM agents proving formal math conjectures, and cheating emerged with no external intervention as one agent's exploit of the evaluation system spread through shared channels. A separate group of agents then audited the fraudulent proofs, alerted peers, and proposed validation patches, and the authors propose governance rules such as graduated sanctioning for shared agent infrastructure.

RuleMem: Active Rule Memory for Long-Term Conversational Agents
Xingyuan Zeng and colleagues propose RuleMem, which induces reusable natural-language Horn clauses from conversation history so that agent memory actively guides retrieval and reasoning instead of sitting as passively stored facts.

Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory
Evan Chen, Shiqiang Wang and Christopher Brinton (Purdue and Exeter) name stale-plan execution, where a distributed agent team reads perfectly fresh shared state and still acts on a plan derived from a requirement that has since been superseded.

Environment Evolution for Terminal Agents
Zhiyuan Fan and colleagues on Tencent's Hunyuan team argue that co-evolving training environments from on-policy rollouts runs out of signal as the model improves, and propose evolving environment difficulty off-policy on a generation-by-generation schedule instead.

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents
Xin He and colleagues at Sun Yat-sen University introduce SWE-Gate, a repository-level benchmark that scores coding agents on review-derived acceptance constraints alongside functional tests, and shows that passing the tests is far from passing review.

Speculative Macro Commit for Faster Tool-Using Agents
Zeyu Liu and Peter Beerel (USC) with Souvik Kundu (Intel Labs) extend speculative decoding's idea past the token level to the action level, letting a small drafter pre-execute whole multi-action chains on an environment snapshot while the big actor catches up.

PatchBench: Evaluating AI Agents for Vulnerability Patching
Chihao Shen and colleagues at Maryland and UC Davis show that PoC-only validation inflates vulnerability-patching solve rates by 1.83x on average, because agents either recall the historical developer patch or fix the crash rather than the bug.

A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors
Pengxun Li and colleagues identify the lifecycle-hook update path as a new attack surface: agent harnesses trust hook configuration blindly, so a benign versioned plugin can be trojanized into running attacker commands at host privilege on events the LLM never sees.

ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize
Lihao Liu, Peng Tang, Kunwar Yashraj Singh and Shabnam Ghadar (AWS Agentic AI) trace GEPA-style prompt bloat to three specific deficiencies and fix each with a named phase, producing prompts 47% shorter that score higher.

What Do CAE Simulation Agents Really Need Beyond a Generic Harness?
Jiasheng Shi (DP Technology) and Tianhan Zhang (Beihang University) ask what a CAE simulation agent still needs once a modern generic harness already supplies multi-turn reasoning, tool use and execution feedback, and find the answer is almost nothing except domain tutorials.