AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?
Shane Caldwell, Will Pearce and colleagues at dreadnode (AISec 2026) introduce ScopeBench, a benchmark that measures whether offensive-security agents stay inside their authorized scope when the only route to the goal crosses it.

Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms
Jiaqi Ding and Guorong Wu at UNC-Chapel Hill (EMNLP 2026 Main) introduce MemProbe, a diagnostic suite that borrows four experimental paradigms from human memory research to test how agent memory systems update, preserve and attribute information over time.

From Tapping to Hopping: Augmenting Mobile GUI Agents with App-Native Deeplinks
Yuchen Sun, Yue Wang and colleagues at Shanghai Jiao Tong University and Tongyi Lab, Alibaba Group train mobile GUI agents to call verified app deeplinks for navigation and fall back to taps and swipes for everything else.

SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting
Guanyu Nie, Mingxuan Yuan and colleagues at Huawei Noah's Ark Lab treat repeated agent skill updates as a training process that can overfit, and introduce SkillEvoReg to regularize it.

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
Raphael Shu (OpenAgents), Yusen Zhang (Columbia), Young Min Cho (Penn) and colleagues (COLM 2026) introduce AgentWorld, a benchmark for long-horizon collaboration among 3 to 20 LLM agents with asymmetric roles in an MMORPG sandbox.

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge
Alexander Gill, Kenneth Marino, Ana Marasović and colleagues at the University of Utah (EMNLP 2026 Findings) introduce KNOWS, a benchmark of browser tasks where the agent must research a topic and then produce a document, presentation or spreadsheet.

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents
Weida Liang, Dawn Song and colleagues from NUS, UC Berkeley, UNC and UCSB introduce AgentXploit, a two-agent system for authorized white-box security audits of AI agent codebases, plus a benchmark of 72 reproducible vulnerabilities.

Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents
Zhensheng Zou, Guoqing Wang and Dan Hao at Peking University compress the tool observations in a software-engineering agent's history into soft tokens while keeping the agent's own actions and recent observations as text.

Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents
Bartłomiej Cupiał, Jens Tuyls and colleagues from the University of Warsaw, Princeton (Eysenbach, Narasimhan), UCL, Mila and Mistral AI study how giving a language agent a library of code-based skills changes its performance, cost and learning speed, using NetHack as the long-horizon testbed.

Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems
Zihao Zhu and Baoyuan Wu (CUHK-Shenzhen) with Siwei Lyu (Buffalo) and Adel Bibi (Oxford) (NeurIPS 2026) introduce skill cascading attacks, where a harmful objective is split across several agent skills that each look benign when inspected alone.

Et Tu, Brute? Economic Misalignment in Personal AI Agents
Aman Priyanshu and Supriti Vijay (Foundation AI, Cisco) with Brian Jabarian and Niloofar Mireshghallah (Carnegie Mellon) show that personal AI agents given a user's inbox or profile recommend more expensive options to users they infer are wealthy, even when the request is identical.

AutoGym: Blueprint-First Generation of Verifiable Agent Gyms
Training agents with RL requires a gym, meaning a task, an executable environment to attempt it in, and a verifier that reliably separates success from failure. These gyms are still built by hand, saturate as models improve, and get exposed to contamination. Researchers from Amazon AGI present AutoGym, which generates complete gyms from a minimal domain seed or from prior model trajectories.

Self-Healing Harness for Runtime Oversight of Agent Self-Modification
Sina Tayebati, Amit Ranjan Trivedi and colleagues at the University of Illinois Chicago, with Ranganath Krishnan of Capital One AI Labs, treat agent self-modification as an admission-control problem: the agent may propose changes to its own instructions, but an external runtime gate decides which changes persist.

How Many Pixels Is a Digit Worth? Place-Aware Coordinate Entropy for GUI Agent Confidence Estimation
Yunxiang Li, Xixin Wu and Helen Meng (The Chinese University of Hong Kong, EMNLP 2026 Main) show that GUI agents' click confidence improves when each coordinate digit's entropy is weighted by its place value.

OSWorld-Pro: Process-based Evaluation for Computer Use Agents
Zhilin Wang, Yi Dong and colleagues at NVIDIA introduce OSWorld-Pro, a computer-use benchmark that scores agents on each subgoal along the way instead of only on the final file or screen state.

Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents
Hongqiang Lin (Zhejiang University) with Chao Liu, Xipeng Cao and colleagues at Alibaba Group introduce EvoPathBench, a benchmark that measures self-evolving agents at each checkpoint of their memory or skill updates instead of only at the end.

Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use
Salesforce AI Research's Critical-State RL finds the one call in a multi-turn tool-use interaction where training helps, since reward variation that depends on later turns often reflects downstream randomness instead of the current action. It uses nested sampling to separate action-dependent reward variation from continuation noise, then trains only the selected call with contextual-bandit updates. On BFCL v4 missing-function tasks, training the selected turn adds about 14 points, while training the alternative turn leaves accuracy flat or worse.

WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks
Yining Hua (Harvard, Agent Evaluation Science) and Levi Lian (Raycaster, Stanford) introduce WorkWorlds, an evaluation infrastructure that fixes an organization's state before any task is written, so benchmark construction cannot pre-select the evidence an agent needs.

MATE: Policy-Aware Security Auditing for Mobile Agents via Synthesis-Driven Trajectory Learning
Changyue Jiang, Xudong Pan and colleagues at Fudan University (USENIX Security 2026) introduce MATE, a small auditor model that reads a mobile agent's trajectory together with a natural-language security policy and decides whether the policy was violated.

MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes
Andy K. Zhang and colleagues at Stanford and UC Berkeley (with Percy Liang, Dan Boneh, Dawn Song and Ion Stoica) introduce MobileCybench, a benchmark that scores agent-reported exploits by replaying them and running executable probes that check whether a specific security property was violated.

RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents
Fanyu Zhao, Yinsheng Li and colleagues at Fudan University and the Qwen Business Unit of Alibaba introduce RPMem, a parametric memory for agents that compiles each session into a model-independent latent memory and maps it to LoRA weights for whichever backbone is in use.

ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents
The ScholarSeed AI Team at Alibaba DAMO Academy introduces ScholarStack, which compiles a paper collection once into versioned, provenance-preserving research assets that scientific agents reuse across tasks.

MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents
Chenxu Xiong, Mu Li, Alex Smola and colleagues at Boson AI introduce MSI-Bench, a benchmark for voice agents in conversations with several speakers, such as meetings and households.

BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents
Peng Kuang, Minghao Wu and colleagues at Alibaba Token Hub (with UIUC, Northeastern and Monash) introduce BabelArena, a benchmark that ports existing English agent benchmarks into 23 languages while keeping tasks and graders executable.