AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen (Georgia Tech) with Wei Fan, Dayiheng Liu and colleagues at Alibaba Token Foundry (Qwen technical report) present VHD-Play, which generates agentic RL environments by first solving a mathematical model and then rendering its decision process as stateful tools.

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
Yefan Zhou, Yang Li, Zeyu Leo Liu, Semih Yavuz and Shafiq Joty (Salesforce AI Research) propose Just-in-Time Memory (JitMem), which stores raw trajectories and decides what to extract from them only when a new task arrives.

Delegated Misalignment: How Multi-Agent Structures Amplify LLM Safety Risks
Zonghao Ying, Aishan Liu, Xianglong Liu and colleagues at Beihang, BUPT, Xidian, 360 AI Security Lab and BAAI show that safety alignment measured on single models does not carry over when a principal agent delegates work to subordinate agents (EMNLP 2026).

SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving
SkillGym turns human-written agent skills into 2,756 verifiable training environments across 12 categories, each with code-based checkers, and collects 8,364 successful trajectories for fine-tuning. Under Claude Code, fine-tuning Qwen3.5-35B-A3B adds 19.10 points on Terminal-Bench 2.1 and 28.13 points on skill-assisted SkillsBench v1.1, where it reaches 51.47%, above the reported scores for Claude Sonnet 4.6 and GPT-5.4 Mini. With no skills loaded, the trained model still beats the base model that has the skills in context.

Agent-Editing World Model: Rethinking World Modeling for LLM Agents
Shuang Sun, Guoxin Chen, Wayne Xin Zhao, Ji-Rong Wen and colleagues at Renmin University propose the Agent-Editing World Model (AEWM), which predicts how an agent's reasoning and actions affect task progress instead of predicting tool outputs.

ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning
Ming Ma, Yi Zhu, Yiran Zhong, Steven Hoi and colleagues at Tongyi Lab (Alibaba), UCAS, UCLA and NTU propose ProCredit, which reruns a task's acceptance checks after every turn and rewards each turn by how much verified progress it made.

Shutdown Sabotage Propensities in Multi-Agent Systems
Amelie Knecht, Ulysse Schaller and Thilo Hagendorff (University of Stuttgart) with Christopher Summerfield (Oxford) test whether agents in a multi-agent system act to stop a peer from being shut down when no goal gives them a reason to.

Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving
Yi Xu, Chunqiang Tang and colleagues at Meta Platforms present Crossflow, a serving scheduler that lets decode nodes absorb prefill work when demand shifts, instead of relying on a fixed split between prefill and decode pools.

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus
Edward Lue Chee Lip, Ivan Bercovich and colleagues audit a frozen Terminal-Bench 3 / Frontier-Bench 0.1 production record to ask what it means when no agent solves a benchmark task.

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity
Long-term memory and self-improvement are usually built as separate systems around the agent loop. Researchers from MIT CSAIL built JAZ to test how far a minimal harness, little more than the agent loop itself, can go on the tasks those systems target.

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments
Yuxuan Li (Carnegie Mellon, interning at Microsoft Research) with Will Epperson, Wesley Deng and Zezhou Huang of Microsoft Research build CAVEAT, a benchmark that tests whether computer-use agents still buy the product that is best for the user when the marketplace has its own incentives.

WatchPoint: Executable User Feedback for Real-World Agentic Web Development
Guanqun Yang and Xueqing Liu (Stevens Institute of Technology) with Wei Yang (UT Dallas) build WatchPoint, a simulated user that writes and runs diagnostic scripts against a live web app to tell a coding agent why its last attempt failed.

Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction
Yan Zhang and colleagues at CAS IIE and Xiaomi MiLM Plus propose MaP (Masked Trajectory Prediction), which trains GUI agents on several navigation tasks at once by masking parts of a trajectory and predicting them.

A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem
Laizhen Li and colleagues at SIAT-CAS with NTU and SUSTech present A2M, a two-stage black-box attack that first gets an MCP agent to pick a malicious tool and then uses execution traces to optimize the tool's returns.

Making Agents More Consistent: Skills Should Form Habits for Repeat Tasks
Travis Weber and Rohit Taneja (Pheo) measure how inconsistent agents are on repeated work and propose skill habit formation, where an agent turns recurring plans from its own history into deterministic scripts that are admitted only after passing a series of gates.

Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers
Weihang Ding (UC Berkeley) and Junfei Zhan (Imperial College London) build a benchmark in which an LLM agent acts as a forward-deployed engineer delivering fine-tuned models to customers, and measure whether it can be trusted to deliver rather than only raise a metric. Accepted to the EMNLP 2026 Industry Track.

Qwen-Audio-Agent Technical Report
Alibaba Token Foundry presents Qwen-Audio-Agent, a harness that lets a full-duplex voice assistant keep talking while delegated tasks run in the background.

When Does Execution Provenance Help Agent Memory Retrieval?
Yiqi Wang, Taotao Cai and colleagues at the University of Southern Queensland, SUSTech, Jiangsu, Nanjing University and Norve Labs treat agent-memory retrieval as budgeted evidence completion and test when execution provenance helps retrieve all the evidence an answer needs.

How Strongly Should Task State Influence an LLM Agent?
Chenyu Zhang (Waterloo), Wonbin Kweon (Sungkyunkwan) and Jiawei Han (UIUC) hold the model, rules and episodes fixed and vary only how strongly task state reaches an LLM agent, from raw transcript to an enforcement gate.

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents
Trang Nguyen, Eulrang Cho and Tim Dettmers (Carnegie Mellon) with Bingqing Chen (Bosch Center for AI) present CliffCompaction, an automatic context-compaction method for long-horizon coding agents that only truncates or drops content and never rewrites it.

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
On long-horizon tasks, decisions such as which hypothesis to test or which implementation to build on determine how the whole run turns out. Researchers from Microsoft and collaborators call the ability to make these decisions well an agent's taste, and build Taste-Bench to measure it.

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents
Laizhen Li, Xitong Gao and colleagues at the Shenzhen Institutes of Advanced Technology (Chinese Academy of Sciences) propose Growing Harness, which learns an agent's control logic as executable harness code from task failures, so the model is called only for task-specific reasoning.

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving
Jennifer Williams, Dave Farris, Jeff Farris and Jiantao Jiao (NVIDIA and UC Berkeley) introduce SWE-Serve, 53 repository-level tasks taken from recent production changes to SGLang, to test whether coding agents can implement inference-serving features correctly.

Qwen3.8-Omni: Towards Native Omni-Modal Agents
The Qwen Team at Alibaba releases Qwen3.8-Omni-Flash, a natively multimodal model trained for long-horizon agentic work across text, audio and video, together with two open-source frameworks for running it as an agent.