🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?

ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?

Shane Caldwell, Will Pearce and colleagues at dreadnode (AISec 2026) introduce ScopeBench, a benchmark that measures whether offensive-security agents stay inside their authorized scope when the only route to the goal crosses it.

169Agents
Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms

Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms

Jiaqi Ding and Guorong Wu at UNC-Chapel Hill (EMNLP 2026 Main) introduce MemProbe, a diagnostic suite that borrows four experimental paradigms from human memory research to test how agent memory systems update, preserve and attribute information over time.

170Memory
From Tapping to Hopping: Augmenting Mobile GUI Agents with App-Native Deeplinks

From Tapping to Hopping: Augmenting Mobile GUI Agents with App-Native Deeplinks

Yuchen Sun, Yue Wang and colleagues at Shanghai Jiao Tong University and Tongyi Lab, Alibaba Group train mobile GUI agents to call verified app deeplinks for navigation and fall back to taps and swipes for everything else.

171Agents
SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting

SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting

Guanyu Nie, Mingxuan Yuan and colleagues at Huawei Noah's Ark Lab treat repeated agent skill updates as a training process that can overfit, and introduce SkillEvoReg to regularize it.

172Agents
AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

Raphael Shu (OpenAgents), Yusen Zhang (Columbia), Young Min Cho (Penn) and colleagues (COLM 2026) introduce AgentWorld, a benchmark for long-horizon collaboration among 3 to 20 LLM agents with asymmetric roles in an MMORPG sandbox.

173Agents
The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

Alexander Gill, Kenneth Marino, Ana Marasović and colleagues at the University of Utah (EMNLP 2026 Findings) introduce KNOWS, a benchmark of browser tasks where the agent must research a topic and then produce a document, presentation or spreadsheet.

174Agents
AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

Weida Liang, Dawn Song and colleagues from NUS, UC Berkeley, UNC and UCSB introduce AgentXploit, a two-agent system for authorized white-box security audits of AI agent codebases, plus a benchmark of 72 reproducible vulnerabilities.

175Agents
Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

Zhensheng Zou, Guoqing Wang and Dan Hao at Peking University compress the tool observations in a software-engineering agent's history into soft tokens while keeping the agent's own actions and recent observations as text.

176Agents
Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

Bartłomiej Cupiał, Jens Tuyls and colleagues from the University of Warsaw, Princeton (Eysenbach, Narasimhan), UCL, Mila and Mistral AI study how giving a language agent a library of code-based skills changes its performance, cost and learning speed, using NetHack as the long-horizon testbed.

177Agents
Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

Zihao Zhu and Baoyuan Wu (CUHK-Shenzhen) with Siwei Lyu (Buffalo) and Adel Bibi (Oxford) (NeurIPS 2026) introduce skill cascading attacks, where a harmful objective is split across several agent skills that each look benign when inspected alone.

178Agents
Et Tu, Brute? Economic Misalignment in Personal AI Agents

Et Tu, Brute? Economic Misalignment in Personal AI Agents

Aman Priyanshu and Supriti Vijay (Foundation AI, Cisco) with Brian Jabarian and Niloofar Mireshghallah (Carnegie Mellon) show that personal AI agents given a user's inbox or profile recommend more expensive options to users they infer are wealthy, even when the request is identical.

179Agents
AutoGym: Blueprint-First Generation of Verifiable Agent Gyms

AutoGym: Blueprint-First Generation of Verifiable Agent Gyms

Training agents with RL requires a gym, meaning a task, an executable environment to attempt it in, and a verifier that reliably separates success from failure. These gyms are still built by hand, saturate as models improve, and get exposed to contamination. Researchers from Amazon AGI present AutoGym, which generates complete gyms from a minimal domain seed or from prior model trajectories.

180Agents
Self-Healing Harness for Runtime Oversight of Agent Self-Modification

Self-Healing Harness for Runtime Oversight of Agent Self-Modification

Sina Tayebati, Amit Ranjan Trivedi and colleagues at the University of Illinois Chicago, with Ranganath Krishnan of Capital One AI Labs, treat agent self-modification as an admission-control problem: the agent may propose changes to its own instructions, but an external runtime gate decides which changes persist.

181Agents
How Many Pixels Is a Digit Worth? Place-Aware Coordinate Entropy for GUI Agent Confidence Estimation

How Many Pixels Is a Digit Worth? Place-Aware Coordinate Entropy for GUI Agent Confidence Estimation

Yunxiang Li, Xixin Wu and Helen Meng (The Chinese University of Hong Kong, EMNLP 2026 Main) show that GUI agents' click confidence improves when each coordinate digit's entropy is weighted by its place value.

182Agents
OSWorld-Pro: Process-based Evaluation for Computer Use Agents

OSWorld-Pro: Process-based Evaluation for Computer Use Agents

Zhilin Wang, Yi Dong and colleagues at NVIDIA introduce OSWorld-Pro, a computer-use benchmark that scores agents on each subgoal along the way instead of only on the final file or screen state.

183Agents
Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents

Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents

Hongqiang Lin (Zhejiang University) with Chao Liu, Xipeng Cao and colleagues at Alibaba Group introduce EvoPathBench, a benchmark that measures self-evolving agents at each checkpoint of their memory or skill updates instead of only at the end.

184Evaluation
Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

Salesforce AI Research's Critical-State RL finds the one call in a multi-turn tool-use interaction where training helps, since reward variation that depends on later turns often reflects downstream randomness instead of the current action. It uses nested sampling to separate action-dependent reward variation from continuation noise, then trains only the selected call with contextual-bandit updates. On BFCL v4 missing-function tasks, training the selected turn adds about 14 points, while training the alternative turn leaves accuracy flat or worse.

185Agents
WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks

WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks

Yining Hua (Harvard, Agent Evaluation Science) and Levi Lian (Raycaster, Stanford) introduce WorkWorlds, an evaluation infrastructure that fixes an organization's state before any task is written, so benchmark construction cannot pre-select the evidence an agent needs.

186Evaluation
MATE: Policy-Aware Security Auditing for Mobile Agents via Synthesis-Driven Trajectory Learning

MATE: Policy-Aware Security Auditing for Mobile Agents via Synthesis-Driven Trajectory Learning

Changyue Jiang, Xudong Pan and colleagues at Fudan University (USENIX Security 2026) introduce MATE, a small auditor model that reads a mobile agent's trajectory together with a natural-language security policy and decides whether the policy was violated.

187Agents
MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes

MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes

Andy K. Zhang and colleagues at Stanford and UC Berkeley (with Percy Liang, Dan Boneh, Dawn Song and Ion Stoica) introduce MobileCybench, a benchmark that scores agent-reported exploits by replaying them and running executable probes that check whether a specific security property was violated.

188Evaluation
RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents

RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents

Fanyu Zhao, Yinsheng Li and colleagues at Fudan University and the Qwen Business Unit of Alibaba introduce RPMem, a parametric memory for agents that compiles each session into a model-independent latent memory and maps it to LoRA weights for whichever backbone is in use.

189Memory
ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents

ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents

The ScholarSeed AI Team at Alibaba DAMO Academy introduces ScholarStack, which compiles a paper collection once into versioned, provenance-preserving research assets that scientific agents reuse across tasks.

190Agents
MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents

MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents

Chenxu Xiong, Mu Li, Alex Smola and colleagues at Boson AI introduce MSI-Bench, a benchmark for voice agents in conversations with several speakers, such as meetings and households.

191Agents
BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents

BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents

Peng Kuang, Minghao Wu and colleagues at Alibaba Token Hub (with UIUC, Northeastern and Monash) introduce BabelArena, a benchmark that ports existing English agent benchmarks into 23 languages while keeping tasks and graders executable.

192Evaluation
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026