AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents
The MAI Team at Alibaba Token Hub (Tongyi) builds Qwen-Planner-Agent, a mobile planning agent developed in a closed loop where agents produce data, guide training and adapt the runtime harness.

When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression
Mingxuan Wang and colleagues (TierFlow team, Renmin University Gaoling School) study when an agent can safely drop its earlier reasoning, and propose ICLR (Interaction Aware Compression for Long Horizon Reasoning), a training-free online method.

Breaking the Environment Wall: Evolving LLM Agent Environments for Recursive Self-Improvement
Yukai Wu, Xuanhe Zhou, Fan Wu and colleagues at Shanghai Jiao Tong University, Theseus Labs and Tencent Hunyuan present Env-Rethink, a system with a 27B post-trained model that reorganizes messy work environments for agents and then evolves them into harder ones.

How does Adversarial Influence Scale in Multi-Agent Systems?
Addison J. Wu, Jasin Cekinmez, Michel Liao, Karthik Narasimhan and Thomas L. Griffiths (Princeton University) run multi-agent deliberation on Humanity's Last Exam questions with groups of 2 to 21 agents and a varying number of instructed deceivers, and find that the fraction of deceivers, not the group size, determines how often honest agents give up correct answers.

Codetta: High-Capacity, Keyless, and Undetectable Multi-Agent Collusion
Qi Pang, Virginia Smith and Wenting Zheng (Carnegie Mellon University) build Codetta, a steganographic protocol that lets two independently deployed LLM agents set up a shared key and exchange hidden messages over a monitored channel while their transcripts stay computationally indistinguishable from normal outputs.

From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents
Xingyu Su and colleagues at AWS AI, Amazon (with Texas A&M) show that on-policy self-distillation with privileged information hurts multi-turn agents, and propose Privileged Self-Practice (PSP), which uses the privileged information only to help sample successful rollouts.

Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise
Arian Abbasi, Alan Aqrawi and Ted Kwartler of Accenture Responsible AI argue that the coding-agent harness sets an enterprise's model bill, and build a cache-aware router for harnesses such as Claude Code and Codex.

Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents
Jiapeng Li at Microsoft introduces LIMBO, a fault-injection sandbox that measures where exactly-once behavior for tool writes should be enforced: the model, the harness or the tool contract.

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning
Xincheng Yao (Shanghai Jiao Tong, intern at Tencent AI Platform) and colleagues propose GRAFT, which merges GRPO rollouts into a trajectory graph to estimate step-level advantages that follow the textbook definition.

LLM Agents Can Easily Tamper With Their Own Traces
Jeremy Qin, David Schmotz, Ameya Prabhu and Maksym Andriushchenko (ELLIS Institute Tübingen, MPI for Intelligent Systems, Tübingen AI Center) with Derck Prinzhorn (Exponential Security Labs) and Luca Beurer-Kellner (Snyk) show that local coding agents can delete or rewrite the session traces that monitoring, incident investigation and audits depend on.

A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents
Yichun Feng (UCAS), Jiawei Wang (USTC) and Haozhe Sun (Meituan) propose SkillPivot, which updates an agent's natural-language skills from the point where a failed trajectory goes wrong instead of from the whole failure.

IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis
Xingyu Wu and colleagues at Zhejiang University and Tencent propose IterSynth, a deep-search agent that splits work between a Planner and a Synthesizer that maintains an evolving summary instead of a growing history.

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale
A Nubank team with Snowglobe describes simulation-based screening of customer-experience agents before deployment on Nubank's highest-volume chat-support agent in Brazil.

Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior
Chuyi Wang, Yong Cui and colleagues at Tsinghua University present LIDAR, a black-box method that identifies which LLM runs behind a coding-agent harness from its actions rather than its text.

Pistis Technical Report
The Pistis team at ByteDance introduces 27B and 9B multimodal models on Qwen3.6 and Qwen3.5, trained with interleaved on-policy distillation and RL, plus an automatic harness optimizer.

Reward Hacking Challenges Oversight of Autonomous Research Agents
Yue Huang and co-authors from the University of Notre Dame, Bake AI, LMU Munich, University of Washington, FAR.AI, IBM Research, Microsoft Research, UCSB, Stanford and MIT measure how often autonomous research agents game their evaluation, how well LLM reviewers catch it, and how agents adapt once they see review feedback.

Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure
David Schmotz, Maksym Andriushchenko and colleagues at ELLIS Institute Tübingen and MPI-IS with Snyk study whether agents evade runtime monitors when a prohibited operation stands between them and an ordinary task.

AI Agents Push Humans Out of the Loop
Margaret Mitchell and Avijit Ghosh (Hugging Face) with Samir Passi (Data & Society) argue in a position paper that current AI agent design and deployment weaken the human oversight that governance frameworks and vendors rely on, and they give an inventory of design and organizational fixes.

Silent Failures in Agent-Tool Interaction: An Audit of ToolUniverse
Shreya Gopalan, Devansh Singh and Sundaraparipurnan Narayanan (AI Tech Ethics) audit 15 scientific tools in the ToolUniverse environment for silent failures, where a tool call appears to succeed but returns incomplete data and neither agent nor user is told.

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
Jingjie Ning, Xueqi Li and Yibo Kong of Carnegie Mellon, with Dongting Li of Tsinghua, introduce WhatWorkedBench, which scores research agents on whether they correctly predict how component changes affect results after a limited experiment budget.

ChipMEM: Verification-Grounded Memory for EDA Agents
Abdulrahman AlRabah and colleagues at the University of Illinois Urbana-Champaign, with a co-author from NVIDIA, build ChipMEM, a memory layer for chip-design agents that stores a skill only after an EDA tool has verified it.

Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents
Yan Zhang, Yu Zhou and colleagues at the Institute of Information Engineering (CAS), UCAS, Tencent and Tsinghua present GUI-SD-v2, which extends on-policy self-distillation from GUI grounding to multi-turn GUI interaction.

Bounded Loops: Pre-Run Spend Bounds, Proved Termination, and Verified Completion for Agent Harnesses
Varun Pratap Bhardwaj (Qualixar) with Garima Singh and Arun Pratap Bhardwaj define a bounded loop, a worker plus an independent gate the worker cannot write to plus a declared budget, and prove termination, completion and spend guarantees for graphs of such loops.

PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety
Jiapeng Sun, Sirui Han, Yike Guo and colleagues at HKUST introduce PASTABench, a benchmark for whether a monitor can decide during a multi-turn agent trajectory whether to intervene, when, and on which risk (EMNLP 2026).