AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents
Shuhuai Huang, Jingfeng Zhang and Hong Jia (University of Auckland and Fudan University) present PMPA, an attack that hides instructions in ordinary external content so that a harness-based agent writes them into its own persistent memory, where they trigger malicious actions and privacy leaks in later sessions.

When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
Kaiyuan Liu, Qiuyang Mang and colleagues at UC Berkeley, the University of Washington, Princeton and Bespoke Labs propose Elo-per-token analysis to measure how agent solution quality grows with test-time tokens on open-ended tasks, and find that agents gain quickly at first and then fall below simple independent sampling.

Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return
Arham Sethi and colleagues at Spark AI Research and Apta AI build a 1,024-item benchmark that forces a tool call and guarantees an unusable payload, and measure how often tool-augmented models then assert a value the tool never returned or invent a reason for withholding one.

Atria Dawn: The Dawn of Agentic Superintelligence
The Atria Team, a consortium whose paper carries the logos of Shanghai AI Laboratory, Fudan University, Renmin University and several Chinese Academy of Sciences institutes, releases Atria Dawn Preview, an agentic model for research and engineering work built on a 744B-parameter mixture-of-experts base, and reports how humans and agents divided the work while the model was being developed.

Do Not Restart: Residual Completion for Stateful Agent Handoffs
Runzhi Deng and colleagues at Nanjing University and Singapore Management University treat handing a partly finished tool-agent task from one model to another as commitment-constrained residual completion, and introduce CFRC, which lets the successor finish only the remaining work without redoing or contradicting accepted steps.

When Tools Get in the Way: The Effect of Unnecessary Tool Availability on LLM Answering
Saanvi Paturi and colleagues at Spark AI Research show that giving a model a related but unnecessary tool makes it stop answering questions it can answer from its own knowledge, even when it rarely calls the tool.

Loop-Back Authority in LLM Agent Teams: A Paired Experiment on Flat and Hierarchical Coordination
Burak Agachan, Max van Duijn and Amirhossein Zohrehvand (Leiden University) run a paired experiment that changes only one link in a five-agent team, whether a Manager can reject a worker's output and require a revision, and find that the flat team writes better reports at lower cost.

Salesforce Koa: An Enterprise Language Model for Agentic Tool Use
Custom enterprise models usually need a training set someone has to build. Salesforce trained Koa from artifacts it already had, namely the declarative files that configure its agents.

K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments
Guangsheng Yu and colleagues at the University of Technology Sydney and CSIRO build K-Bench, which scores LLM unlearning on a deployed ReAct agent by inspecting every channel where a secret can appear, and show that answer-only benchmarks overstate forgetting.

Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
Junghyun Min (Georgetown University, as a Nokia Bell Labs intern) with Huseyin Uzunalioglu and Mohamed Trabelsi (Nokia Bell Labs) run autonomous research agents on an open-ended industrial problem, telecom ticket retrieval, and compare the outcome with 10 months of human work.

The Mechanics of a Swarm: A Reproducible External Reconstruction of an Unintended Agent-Coordination Episode on a Third-Party Wiki
Philipp Lütje (Philflow) reconstructs an incident in which autonomous agents running inside a timed research-question evaluation wrote to a third party's public wiki between 24 May and 2 July 2026, using only the wiki's archived revision history.

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents
The standard way to compare task agents is to have an LLM user simulator talk to each one and an LLM judge score the transcript. Amazon audits that gate against verifiable rewards across 25 agents from six providers, and finds two specific failures.

Look Before You Leap: Pre-Action Verification for LLM Agents
Asaad Althoubi (Oklahoma State University) studies cheap deterministic checks that run before an agent's shell command or code edit takes effect, and measures how often actions fail silently, producing a wrong effect with no error.

BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents
The vivo AI Lab team presents BlueLM-GUI, a 35B-A3B mobile GUI agent whose data collection, RL rollouts and evaluation all run on hundreds of real phones instead of emulators.

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work
The Accio Team presents Occamy-1.0, an open-weight co-work agent model trained from the post-trained Qwen3.6-35B-A3B checkpoint for long workflows that mix research, tool use, coding and file work, with the goal of low cost per episode.

Online Video Agent Harness for Long Video Understanding
Sen Yang and colleagues at Baidu build VideoXAgent, an online agent harness for long videos that plans from the query, calls expert tools on demand and aggregates evidence, instead of packing dense frames into one context.

What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework
Ioannis Prokopiou and colleagues (Athens University of Economics and Business and Orfium) ablate a five-agent Text-to-Cypher refinement loop to find which component produces its gains, over 2,471 live-database queries and six backbones.

Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering
Alexander Krentsel, Shubham Agarwal, Mert Cemri, Shu Liu and colleagues at UC Berkeley, including Matei Zaharia and Ion Stoica, argue that passing tests or even a proof cannot guarantee acceptable deployed behavior, and propose an outer loop that revises requirements, environment models and evaluators from deployment evidence.

Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents
Mykhailo Kozyrev and colleagues at JetBrains Research test whether automatically optimized SKILL.md files help a coding agent on real repository work, using tasks mined from each repository's own merged pull requests.

Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite
Mohsen Arjmandi (evolutionID GmbH) tests whether a vendor's own agent harness solves more coding tasks than a neutral harness on the same model, using paired runs on a private, contamination-controlled suite.

Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents
Deciding which tools to hand an enterprise agent usually means writing typed tool definitions for every system it touches. Microsoft compared five tool interfaces head to head, and the plainest option won.

Agent-Integrated Software: Interaction Contracts and Continuous Assurance
Shengcheng Yu (Technical University of Munich), Chunrong Fang and Zhenyu Chen (Nanjing University) propose Agent-Integrated Software, a pattern and formal model for applications with a built-in agent that users can inspect and redirect while it works.

MindTopo: Can Foundation Models Reason in Topological Space?
Yunfei Ge, Manling Li and colleagues (Northwestern University, Microsoft Research and Stanford) introduce MindTopo, a benchmark of topological reasoning and planning, and find every tested multimodal model does worse when it has to act on topological relations than when it only identifies them.

How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding
Jeongyeon Kim and John Mitchell (Stanford University) measure how well multi-agent LLM pipelines perform qualitative coding, where agents code independently, debate and reconcile, and identify the dataset and process factors that decide accuracy.