🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

Xinjie Shen (Georgia Tech) with Wei Fan, Dayiheng Liu and colleagues at Alibaba Token Foundry (Qwen technical report) present VHD-Play, which generates agentic RL environments by first solving a mathematical model and then rendering its decision process as stateful tools.

217Agents
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

Yefan Zhou, Yang Li, Zeyu Leo Liu, Semih Yavuz and Shafiq Joty (Salesforce AI Research) propose Just-in-Time Memory (JitMem), which stores raw trajectories and decides what to extract from them only when a new task arrives.

218Memory
Delegated Misalignment: How Multi-Agent Structures Amplify LLM Safety Risks

Delegated Misalignment: How Multi-Agent Structures Amplify LLM Safety Risks

Zonghao Ying, Aishan Liu, Xianglong Liu and colleagues at Beihang, BUPT, Xidian, 360 AI Security Lab and BAAI show that safety alignment measured on single models does not carry over when a principal agent delegates work to subordinate agents (EMNLP 2026).

219Safety
SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

SkillGym turns human-written agent skills into 2,756 verifiable training environments across 12 categories, each with code-based checkers, and collects 8,364 successful trajectories for fine-tuning. Under Claude Code, fine-tuning Qwen3.5-35B-A3B adds 19.10 points on Terminal-Bench 2.1 and 28.13 points on skill-assisted SkillsBench v1.1, where it reaches 51.47%, above the reported scores for Claude Sonnet 4.6 and GPT-5.4 Mini. With no skills loaded, the trained model still beats the base model that has the skills in context.

220Agents
Agent-Editing World Model: Rethinking World Modeling for LLM Agents

Agent-Editing World Model: Rethinking World Modeling for LLM Agents

Shuang Sun, Guoxin Chen, Wayne Xin Zhao, Ji-Rong Wen and colleagues at Renmin University propose the Agent-Editing World Model (AEWM), which predicts how an agent's reasoning and actions affect task progress instead of predicting tool outputs.

221Agents
ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning

ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning

Ming Ma, Yi Zhu, Yiran Zhong, Steven Hoi and colleagues at Tongyi Lab (Alibaba), UCAS, UCLA and NTU propose ProCredit, which reruns a task's acceptance checks after every turn and rewards each turn by how much verified progress it made.

222Reinforcement Learning
Shutdown Sabotage Propensities in Multi-Agent Systems

Shutdown Sabotage Propensities in Multi-Agent Systems

Amelie Knecht, Ulysse Schaller and Thilo Hagendorff (University of Stuttgart) with Christopher Summerfield (Oxford) test whether agents in a multi-agent system act to stop a peer from being shut down when no goal gives them a reason to.

223Agents
Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving

Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving

Yi Xu, Chunqiang Tang and colleagues at Meta Platforms present Crossflow, a serving scheduler that lets decode nodes absorb prefill work when demand shifts, instead of relying on a fixed split between prefill and decode pools.

224Agents
What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

Edward Lue Chee Lip, Ivan Bercovich and colleagues audit a frozen Terminal-Bench 3 / Frontier-Bench 0.1 production record to ask what it means when no agent solves a benchmark task.

225Agents
Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

Long-term memory and self-improvement are usually built as separate systems around the agent loop. Researchers from MIT CSAIL built JAZ to test how far a minimal harness, little more than the agent loop itself, can go on the tasks those systems target.

226Agents
CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

Yuxuan Li (Carnegie Mellon, interning at Microsoft Research) with Will Epperson, Wesley Deng and Zezhou Huang of Microsoft Research build CAVEAT, a benchmark that tests whether computer-use agents still buy the product that is best for the user when the marketplace has its own incentives.

227Agents
WatchPoint: Executable User Feedback for Real-World Agentic Web Development

WatchPoint: Executable User Feedback for Real-World Agentic Web Development

Guanqun Yang and Xueqing Liu (Stevens Institute of Technology) with Wei Yang (UT Dallas) build WatchPoint, a simulated user that writes and runs diagnostic scripts against a live web app to tell a coding agent why its last attempt failed.

228Agents
Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction

Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction

Yan Zhang and colleagues at CAS IIE and Xiaomi MiLM Plus propose MaP (Masked Trajectory Prediction), which trains GUI agents on several navigation tasks at once by masking parts of a trajectory and predicting them.

229Agents
A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem

A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem

Laizhen Li and colleagues at SIAT-CAS with NTU and SUSTech present A2M, a two-stage black-box attack that first gets an MCP agent to pick a malicious tool and then uses execution traces to optimize the tool's returns.

230Agents
Making Agents More Consistent: Skills Should Form Habits for Repeat Tasks

Making Agents More Consistent: Skills Should Form Habits for Repeat Tasks

Travis Weber and Rohit Taneja (Pheo) measure how inconsistent agents are on repeated work and propose skill habit formation, where an agent turns recurring plans from its own history into deterministic scripts that are admitted only after passing a series of gates.

231Agents
Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

Weihang Ding (UC Berkeley) and Junfei Zhan (Imperial College London) build a benchmark in which an LLM agent acts as a forward-deployed engineer delivering fine-tuned models to customers, and measure whether it can be trusted to deliver rather than only raise a metric. Accepted to the EMNLP 2026 Industry Track.

232Agents
Qwen-Audio-Agent Technical Report

Qwen-Audio-Agent Technical Report

Alibaba Token Foundry presents Qwen-Audio-Agent, a harness that lets a full-duplex voice assistant keep talking while delegated tasks run in the background.

233Agents
When Does Execution Provenance Help Agent Memory Retrieval?

When Does Execution Provenance Help Agent Memory Retrieval?

Yiqi Wang, Taotao Cai and colleagues at the University of Southern Queensland, SUSTech, Jiangsu, Nanjing University and Norve Labs treat agent-memory retrieval as budgeted evidence completion and test when execution provenance helps retrieve all the evidence an answer needs.

234Retrieval
How Strongly Should Task State Influence an LLM Agent?

How Strongly Should Task State Influence an LLM Agent?

Chenyu Zhang (Waterloo), Wonbin Kweon (Sungkyunkwan) and Jiawei Han (UIUC) hold the model, rules and episodes fixed and vary only how strongly task state reaches an LLM agent, from raw transcript to an enforcement gate.

235Agents
CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

Trang Nguyen, Eulrang Cho and Tim Dettmers (Carnegie Mellon) with Bingqing Chen (Bosch Center for AI) present CliffCompaction, an automatic context-compaction method for long-horizon coding agents that only truncates or drops content and never rewrites it.

236Code
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

On long-horizon tasks, decisions such as which hypothesis to test or which implementation to build on determine how the whole run turns out. Researchers from Microsoft and collaborators call the ability to make these decisions well an agent's taste, and build Taste-Bench to measure it.

237Agents
Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

Laizhen Li, Xitong Gao and colleagues at the Shenzhen Institutes of Advanced Technology (Chinese Academy of Sciences) propose Growing Harness, which learns an agent's control logic as executable harness code from task failures, so the model is called only for task-specific reasoning.

238Agents
SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

Jennifer Williams, Dave Farris, Jeff Farris and Jiantao Jiao (NVIDIA and UC Berkeley) introduce SWE-Serve, 53 repository-level tasks taken from recent production changes to SGLang, to test whether coding agents can implement inference-serving features correctly.

239Evaluation
Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen3.8-Omni: Towards Native Omni-Modal Agents

The Qwen Team at Alibaba releases Qwen3.8-Omni-Flash, a natively multimodal model trained for long-horizon agentic work across text, audio and video, together with two open-source frameworks for running it as an agent.

240Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026