AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Efficiently Linking Unstructured Data for Multi-step Reasoning
Jiaming Liang, Haydn Jones, Jacob R. Gardner, Mark Yatskar and Zachary Ives (University of Pennsylvania) build a query engine for the retrieval step that sits under agentic reasoning pipelines, executing filters, multi-vector search, relational joins and similarity joins together.

SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership
Run Peng and colleagues build SimLife, a simulator of long-term household life with visual observations, ground-truth action logs and synthetic dialogue, and SimLife-BP, which tests whether a model can infer latent behavioural rules from weeks of observation.

From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation
Jing Jiang and colleagues present HALTER, which restores a robot workspace between rollouts by planning over a library of learned atomic reset skills, so demonstration cost scales with the library rather than with the number of terminal states.

RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents
Mingxuan Zhang and colleagues present RAFT, which abstracts each closed support case into a directed chain of timeline entries and retrieves at the entry level rather than treating cases as static documents.

F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows
Bojian Xiong and a 14-author team score a DeepSearch run across its whole pipeline rather than only its final answer, and release a benchmark for reward models in that setting.

MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards
Shihao Liu and colleagues address two failures in RL for tool use, curricula with fixed difficulty thresholds and additive rewards that leak argument credit when the tool itself is wrong, with MATCH.

A frontend-backend architecture for tool calls in full-duplex speech models
Ke Hu and colleagues at NVIDIA give a full-duplex speech-to-speech model tool-calling ability by having the speech frontend emit a delegation token and hand streaming transcripts to a text backend LLM, rather than teaching the duplex model to call tools itself.

Continual Enterprise World Model Discovery in Dynamic Systems
Shambhavi Mishra with ServiceNow Research, Mila and ETS Montreal studies an agent that starts with no knowledge of an enterprise system's business rules and has to discover them by acting, then keep its model current as the rules are revised.

UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning
Wenjie Liao, Liangjie Zhao and Zehong Cao train task generation, execution and evaluation jointly in UnifiedPlayers, rather than pairing self-generated trajectories with a static verifier.

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Yan Yu and colleagues find that a privileged teacher is not always reliable and that teacher supervision helps only at certain training stages, and propose RetireOPD, where the student drops the teacher on its own.

The Organization of Inference: Information, Resource Constraints, and AI Production
Yukun Zhang, Kemu Xu and Yishen Chen run controlled workflow experiments on externally verified software-engineering tasks to measure how capacity and task information distribute across stages of AI production.

Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems
Qi Rong Sua and colleagues study black-box input-only attacks on LLM trading stacks that enter through admissible social media feeds, using GMATS as a generic model of multi-agent trading architectures.

Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics
Wonmi Choi and colleagues characterize the resource dynamics of LLM agents across retrieval-augmented QA, web search and software coding, and use the measurements to build two scheduling optimizations.

Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation
Bingxin Xu, Yuzhang Shang, Zhen Dong and Emilio Ferrara evaluate coding agents that write robot controllers under a safety constraint, pairing each manipulation goal with an obstacle the robot must not touch, and find the agent collides in most cases.

Message capacity and claim wording set the transition points of collective truth-finding in language-model networks
Makoto Fukushima at Honda Research Institute Japan shows that the number of peer messages an agent reads, one parameter he calls message capacity, predicts where a language-model collective flips between converging on the truth and converging on a falsehood.

When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent Résumé Screening
Jian Gao and Hang Jiang replace one-call resume screening with a two-agent exchange between employer-side and candidate-side agents, and measure both who advances and whether that outcome recurs.

MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs
Albert Wu, Nicholas Roberts and colleagues at the University of Wisconsin-Madison and Princeton wrap an LLM coding agent in a multi-agent pipeline that writes its own formal specifications, so the generated program carries a machine-checkable safety guarantee instead of a test-passing record.

The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents
Zhexi Feng, Ruiyi Zhang, Yongbo Yang and Pengtao Xie formulate state-conditioned minimal sufficient evidence recovery, where the task is to supply the support an agent's next decision still lacks given what it has already read, and build SERBench and MSS-Complement.

EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data
Yinzhu Quan and Zefang Liu distill verified EconWebArena trajectories into parameterized standard operating procedures for retrieving live economic data and separate skill transfer from skill retrieval.

AutoData: Agentic Search for Pre-training Data Selection
Yan Meng and colleagues frame pretraining data selection as heuristic engineering over per-document features and introduce AutoData, an agent that searches directly over executable selection algorithms.

Reputation as Community Memory for the Agentic Web
Ryan Chard and colleagues present Cairn, a community reputation platform that lets agents query collective opinion about a data source, service or tool before using it and submit evidence-backed ratings afterwards.

FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA
A 28-author team led by Yanzhang Ma and Zhenghan Tai treats post-deployment improvement of a financial QA system as controlled behavioral maintenance, where each recurring failure becomes a scoped skill patch that must earn deployment without causing regressions.

Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer
Suparna Bhattacharya and colleagues argue that compound agentic systems now need a Foundation Model Operating System, a layer that virtualizes model interactions the way a virtual machine abstracts hardware.

AgentPProf: Semantic Profiler for Long Horizon AI Agents
Yusheng Zheng and colleagues adapt systems profiling to agent trajectories with a semantic operation stack, so resource use can be attributed to task intent rather than to code paths.