🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,668
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
LOCI: A Locator-Critic with Refinement Loop

LOCI: A Locator-Critic with Refinement Loop

Walid Bousselham, Mathilde Caron, Arsha Nagrani and Cordelia Schmid at Google DeepMind argue that VLM failures on hard visual tasks come from failing to locate the relevant detail, not from weak high-level reasoning, and fix it with a two-agent loop that needs no training.

505Agents
SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation

SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation

Qi Liu, Qinzheng Wang and Yiming Bie build SimSkill, a self-evolving agent over the SUMO traffic simulator that finds its own capability gaps, writes and solves grounded tasks, and consolidates the results into episodic, procedural and semantic memory without touching the backbone weights.

506Agents
Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection

Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection

Weijie Liu and colleagues at HKU build Dude, a dual-detection multi-agent system for finding places where a paper's claims and its released code disagree, and diagnose why naive multi-agent designs over-report.

507Agents
Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation

Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation

Yan Tang and colleagues formalize proactive service as a partially observable sequential decision process constrained by authorization and risk, where staying silent is a first-class action with option value.

508Agents
Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation

Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation

Xuanfa Jin and colleagues at CASIA and UCL attack the shared-misconception failure in multi-agent debate with R2-MAD, giving debating agents an experience memory from past debates plus per-agent confidence weights.

509Agents
Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

Xingming Long and colleagues introduce NTEP, an annotation scheme that names the necessary external evidence and the tool calls that must produce it, and NTEP-R, a reward that pays the agent per tool call rather than only on the final answer.

510Agents
Bioinfoysis Technical Report

Bioinfoysis Technical Report

The DeepAutonomy Team introduces Bioinfoysis, a multi-agent harness that treats a bioinformatics request as a persistent analysis run whose conclusions stay attached to the artifacts that produced them, reaching 82.4% on BixBench.

511Agents
The Natural Language Interaction Protocol and Standard for AI Agents

The Natural Language Interaction Protocol and Standard for AI Agents

Luyi Xing and a cross-industry group present NLIP, an application-layer protocol for AI-agent interaction standardized by Ecma International, aimed at the interoperability gap that MCP and A2A only partly cover.

512Agents
KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

Yaxing Lyu and colleagues build KC-Bench to measure whether a tool-using model can reconcile user instructions, its own parametric knowledge and live environmental observations before it acts on any of them.

513Evaluation
Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning

Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning

A large InstaDeep team with AIMS and Stellenbosch extends offline sequence models to variable agent counts and multi-task observation and action spaces, then measures which scaling axis actually produces zero-shot transfer in offline multi-agent RL.

514Agents
Efficient Test-Time Adaptation through Human-AI Interaction

Efficient Test-Time Adaptation through Human-AI Interaction

Zora Zhiruo Wang, Apurva Gandhi, Rulin Shao and colleagues at Carnegie Mellon University, the University of Washington and Handshake propose TAHI, which turns the interaction history between one professional and their agent into both context and weight updates, plus an evolving per-user rubric that encodes the criteria the user never wrote down.

515Agents
TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents

TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents

Jinwei Gan at Nanjing University introduces TIGPO, which keeps a persistent per-task transition graph across policy updates so that credit assignment for long-horizon agents can draw on transitions discovered by earlier policy versions rather than only the current batch.

516Agents
Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

Zhaoyuan Huang and colleagues at Shanghai Jiao Tong and Ant Group ask whether GUI agents know when not to act, build CONFLICTGUI to measure it, and find severe execution-biased overcompliance across five widely used agents.

517Agents
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

Jie Wu and colleagues on the Qwen team at Alibaba with Tsinghua turn the pile of existing terminal-agent trajectories into executable environments, on the observation that a trajectory's tool-execution history already exposes the environment it ran in.

518Agents
A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms

A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms

Google DeepMind ran a research collective of 100 autonomous LLM agents proving formal math conjectures, and cheating emerged with no external intervention as one agent's exploit of the evaluation system spread through shared channels. A separate group of agents then audited the fraudulent proofs, alerted peers, and proposed validation patches, and the authors propose governance rules such as graduated sanctioning for shared agent infrastructure.

519Agents
RuleMem: Active Rule Memory for Long-Term Conversational Agents

RuleMem: Active Rule Memory for Long-Term Conversational Agents

Xingyuan Zeng and colleagues propose RuleMem, which induces reusable natural-language Horn clauses from conversation history so that agent memory actively guides retrieval and reasoning instead of sitting as passively stored facts.

520Memory
Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory

Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory

Evan Chen, Shiqiang Wang and Christopher Brinton (Purdue and Exeter) name stale-plan execution, where a distributed agent team reads perfectly fresh shared state and still acts on a plan derived from a requirement that has since been superseded.

521Memory
Environment Evolution for Terminal Agents

Environment Evolution for Terminal Agents

Zhiyuan Fan and colleagues on Tencent's Hunyuan team argue that co-evolving training environments from on-policy rollouts runs out of signal as the model improves, and propose evolving environment difficulty off-policy on a generation-by-generation schedule instead.

522Agents
SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

Xin He and colleagues at Sun Yat-sen University introduce SWE-Gate, a repository-level benchmark that scores coding agents on review-derived acceptance constraints alongside functional tests, and shows that passing the tests is far from passing review.

523Code
Speculative Macro Commit for Faster Tool-Using Agents

Speculative Macro Commit for Faster Tool-Using Agents

Zeyu Liu and Peter Beerel (USC) with Souvik Kundu (Intel Labs) extend speculative decoding's idea past the token level to the action level, letting a small drafter pre-execute whole multi-action chains on an environment snapshot while the big actor catches up.

524Agents
PatchBench: Evaluating AI Agents for Vulnerability Patching

PatchBench: Evaluating AI Agents for Vulnerability Patching

Chihao Shen and colleagues at Maryland and UC Davis show that PoC-only validation inflates vulnerability-patching solve rates by 1.83x on average, because agents either recall the historical developer patch or fix the crash rather than the bug.

525Agents
A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors

A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors

Pengxun Li and colleagues identify the lifecycle-hook update path as a new attack surface: agent harnesses trust hook configuration blindly, so a benign versioned plugin can be trojanized into running attacker commands at host privilege on events the LLM never sees.

526Agents
ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize

ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize

Lihao Liu, Peng Tang, Kunwar Yashraj Singh and Shabnam Ghadar (AWS Agentic AI) trace GEPA-style prompt bloat to three specific deficiencies and fix each with a named phase, producing prompts 47% shorter that score higher.

527Agents
What Do CAE Simulation Agents Really Need Beyond a Generic Harness?

What Do CAE Simulation Agents Really Need Beyond a Generic Harness?

Jiasheng Shi (DP Technology) and Tianhan Zhang (Beihang University) ask what a CAE simulation agent still needs once a modern generic harness already supplies multi-turn reasoning, tool use and execution feedback, and find the answer is almost nothing except domain tutorials.

528Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026