AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control
Minsun Shim and colleagues at UC Irvine and other University of California campuses show three new attacks that make personal agents leak private data through ordinary interaction, and propose FLOWSEAL, which enforces confidentiality outside the model.

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
Harsh Raj and colleagues at Scale AI treat root-cause attribution of long agent failures as a search problem and introduce Continual Search, which prompts an LLM judge over several turns to keep looking for unresolved evidence instead of settling on its first plausible diagnosis.

Rollback the World, Keep the Reflection: Rollback-Induced Reflection for Long-Horizon LLM Agents
Yi Yu, Liuyi Yao, Yaliang Li and colleagues at Wuhan University and Alibaba Group propose Rollback-Induced Reflection, which restores an agent's environment to an earlier state while keeping lessons from the abandoned trajectory.

ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
Siwei Wu, Chenghua Lin and colleagues at Beihang University, the University of Manchester and collaborating institutions propose ModularRSI, a framework for evolving agent harnesses that transfer to unseen tasks rather than overfitting the benchmark used during evolution.

MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents
Jiang, Yuan and Li build a benchmark that measures rare high-severity memory failures in long-horizon agents per risk category, on the argument that an aggregate accuracy score hides exactly the events that matter.

OpenAI4S: Code as Action, Science as Sessions
Gongbo Zhang, Li Yuan and colleagues at Peking University Shenzhen Graduate School release OpenAI4S, an open-source research agent that runs scientific actions as code cells in persistent Python and R kernels with full provenance tracking.

Coaching Qwen3 Coder 30B to Think Like a CodeClash Arena Agent
Ivy Ning Zhang (Stanford) post-trains Qwen3-Coder-30B on stronger agents' CodeClash trajectories to improve its multi-round arena play.

LIMBO: Lifelong Inference-Time Memory and Budget Optimization for LLM Agents
Siddharth Sharma and colleagues at UC San Diego and West Virginia University introduce LIMBO, an online method that decides for each incoming task how much past experience a lifelong agent should replay into its prompt and how much inference budget to spend.

HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning
Hongliang Wei and colleagues at Harbin Institute of Technology and Alibaba Cloud train one policy across several agent harnesses and introduce HarnessBandit, an online scheduler that picks which harness to train on at each optimizer step.

CoMem: Collective-Individual Memory Synergy for Evolutionary Multi-Agent Systems
Yu and colleagues propose a two-tier memory for multi-agent systems that keeps each agent's private experience separate from the group's shared knowledge, so shared memory does not erase what makes individual agents different.

Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Tong Zheng and colleagues at the University of Maryland and Google DeepMind introduce Dream-RSI, which improves a coding agent's exploration policy by replaying its own past discovery trees as a simulator.

VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets
Yu Bai and colleagues (Zhongguancun Laboratory, Tsinghua University and China Mobile) build VRL-Bench to compare verbal trial-and-error learning methods such as Reflexion under a fixed trial budget, and propose a scheduler that splits the budget between exploiting reflections and exploring.

BusMA: A Bus Communication Substrate for Multi-Agent Systems
Peng, Zhang, Wang and Aletras replace the manager-worker and router topologies used in most multi-agent systems with a shared bus, so any agent can address any peer directly instead of routing through a coordinator.

Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents
Grace Chang Yuan, Pranav Rajpurkar and colleagues at MIT and Harvard Medical School study agents that manage a full emergency-department shift and introduce Asclepius, a harness that rewrites its own operating manual between shifts.

BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents
Rao and Jaggi build a measurement harness that makes the per-call input-token budget the independent variable when comparing agent memory strategies, and report budget-violation rates as a first-class outcome rather than a footnote.

The Stochastic Deputy: Structural Tenant Isolation for Tool-Using LLM Agents
Baig and colleagues formalize what happens when a multi-tenant tool asks an LLM agent to supply the tenant identifier, and show that removing the parameter from the tool schema is a stronger defense than validating it.

RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
Sibo Zhu and colleagues introduce RSIAgent, a training-free multi-agent framework in which curriculum, actor and verifier agents explore a new environment and build a reusable memory of its causal structure.

CovR: Coverage-Aware Hardware Verification via Reasoning-Guided Reinforcement Learning
Abdelatty, Nouh and Reda (Brown University) build CovR, an agentic testbench-generation system for RTL hardware verification that optimizes for coverage rather than functional correctness alone, and distill the resulting behavior into a student model with simulation-derived rewards.

Efficiently Linking Unstructured Data for Multi-step Reasoning
Jiaming Liang, Haydn Jones, Jacob R. Gardner, Mark Yatskar and Zachary Ives (University of Pennsylvania) build a query engine for the retrieval step that sits under agentic reasoning pipelines, executing filters, multi-vector search, relational joins and similarity joins together.

AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation
Keshu Wu and colleagues at Texas A&M and collaborators treat air-ground co-simulation scenario generation as compilation with verification, so a scenario that runs is also checked against the relationships the user asked for.

A Scalable Trust Discovery Architecture for the Internet of Agents
Song Zhang and colleagues propose a three-layer registry and resolver architecture for agent discovery, addressing the part of agent protocols that tool invocation standards leave open.

Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions
Xiaofei Yuan and colleagues compare full replanning, classical plan repair and LLM-based local revision on the same disrupted-itinerary benchmark, which prior work could not do because each method defined the task differently.

PAPC: Platform Mediation for Privacy-Propagation Externalities in AI-Mediated Workflows
Huang, Wu, Hou and Zheng model privacy loss in multi-principal agent platforms as an externality created by intermediate events rather than by the final answer, and build PAPC, a platform layer that intercepts every information-moving event before it reaches shared state.

Language-model groups overstate consensus when replaying human deliberation on a reasoning task
Tengfei Shao replays 100 held-out human Wason group discussions with matched LLM agent groups and finds the agent groups reach full consensus far more often than the people they stand in for.