AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Coding-Agent Benchmarks Should Match Their Users' Task Flows
Igor Slinko, Yaroslav Golubev and Sergey Titov (JetBrains Research) compare 4,782 real coding-agent sessions from JetBrains IDEs with issue-derived benchmarks and propose SWE-TaskFlow to reshape benchmarks toward a measured interaction pattern.

Agent Plasticity: Measuring Self-Improvement Through Experience
Harman Singh, Anirudh Goyal, Jason Weston, Gabriel Synnaeve, Rob Fergus and colleagues at Meta Superintelligence Labs (with UC Berkeley and Princeton's Sanjeev Arora) propose agent plasticity, a measure of how efficiently an agent turns experience into gains on held-out tasks.

Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver
Egor Pakhomov and Erik Nijkamp (Salesforce AI Research) test what MemoryAgentBench's Conflict Resolution split measures by running its own stated rule, newest statement wins, as a zero-learning resolver. Accepted at the NeurIPS 2026 Interpreting Agent Behavior workshop.

Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System
Panagiotis Kasnesis and colleagues (University of West Attica) evaluate 9 models from 0.8B parameters to a hosted frontier model at each of the five LLM call sites of Wactorz, a deployed open-source multi-agent home-automation framework. Accepted at a NeurIPS 2026 workshop on SLMs for agentic systems.

SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles
Yuyao Ge and colleagues at the Institute of Computing Technology, Chinese Academy of Sciences (with UC Merced and Tsinghua) present SkillForge, an agentic RL method in which the skill library and the policy are updated together. Accepted at NeurIPS 2026.

DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents
Jike Zhong, Ritwick Chaudhry, Nishant Sankaran and colleagues at Amazon AGI (with USC) introduce DSV-Mem, a benchmark for dense, stateful visual memory in multimodal agents that assist with professional workflows.

From Evidence to Action: How Tool-Using Agents Fail
Hongzhan Lin, Shidong Cao, Ziyang Luo and Wenhao Chai (Princeton) with Mong-Li Lee and Wynne Hsu (NUS) and authors at HKBU and Amazon Web Services study where tool-using agents break the chain from established evidence to action, and release SafeActBench.

ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?
Haizhong Zheng, Yizhuo Di, Ranajoy Sadhukhan, Shuowei Jin and Beidi Chen (Carnegie Mellon, Infini-AI Lab) introduce ServeLearnBench, a benchmark for agents that must infer and revise hidden environment policies from serving experience.

Structuring MoE Expert Selection for Agentic Reinforcement Learning
Bolian Li (Apple and Purdue) with Ting-Yao Hu, Cheng-Yu Hsieh, Oncel Tuzel, Raviteja Vemulapalli and colleagues at Apple study how mixture-of-experts routing relates to agentic behavior and control it during RL post-training.

Stateless Language Agents: Scaling Long-Horizon Automated Research
Qizheng Zhang, Changxiu Ji, Kunle Olukotun and colleagues at Stanford, with CMU, UW and SambaNova, introduce Stateless Language Agents (SLA), where the harness holds all research state and every agent call starts from a freshly built context.

ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents
Yupeng Su, Jiayi Tian and Zheng Zhang (UC Santa Barbara) with Souvik Kundu (Intel) present ReFold, a training-free rendering layer that compresses what a long-horizon agent sees while keeping the full interaction history recoverable.

When to Remember, When to Abstain: Category-Conditioned Retention for Reliable Agent Memory
Olukunle Owolabi, Pulkit Gupta and Fei Wang at Meta AI study when an agent's memory pipeline should store an inferred assertion, and propose confidence thresholds conditioned on the assertion's semantic category. The paper is accepted at the NeurIPS 2026 Social Agent Workshop.

AMBER: Training Long-Horizon Web Agents through Append-Only Memory
Chinmay Savadikar, Tianfu Wu, Lingyun Wang and colleagues at North Carolina State University and Shopify introduce AMBER, an append-only memory that a web agent learns to write while it reasons and acts, trained end-to-end with outcome-reward RL.

NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale
Songlin Jiang and Mario Di Francesco (Aalto University) with Zhiyu Li, Terry Kong and colleagues at NVIDIA present NeMo-DCR, a bit-exact delta-compressed weight synchronization (refit) for agentic RL, released in NVIDIA NeMo RL.

Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale
Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro and Lorenzo Dini at Google describe FlowAgent, a program repair agent deployed across Google that fixes test failures during pre-submit continuous integration, before the developer moves on to other work. The paper is accepted at ASE 2026.

SquidAgent: Parallelize Wisely, Coordinate Efficiently
Yexiong Lin, Tongliang Liu and colleagues at the University of Sydney with Shanshan Ye (MBZUAI) present SquidAgent, which decides layer by layer whether to parallelize agent work; the paper is accepted at NeurIPS 2026.

Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations
Toby D. Pilditch, Konstantinos Voudouris, Alexandra Abbas and Cozmin Ududec at the UK AI Security Institute release Transect, an open-source package built on Inspect Scout for analysing very long agent evaluation transcripts in a reproducible way.

MemTrace: State-Consistent Memory for Long-Horizon Coding Agents
Hongming Xu, Zhiyu Li, Juncheng Zhang and colleagues at Shanghai Jiao Tong University and MemTensor introduce MemTrace, a memory system for long-horizon coding agents that checks recalled evidence against the current repository state before reusing it.

MLLMs Fail to Refuse when Using Tools Agentically
Rikiya Takehi (MIT, during an internship at NVIDIA) with Ryo Hachiuma, Shaona Ghosh and colleagues at NVIDIA show that giving multimodal LLMs tools makes them refuse harmful requests less often; the paper is accepted at NeurIPS 2026.

Causal Improvement Graph for Agentic Harness Optimization
Junjie Zhang, Shunyu Liu and Dacheng Tao (Nanyang Technological University) with Ting-En Lin and Yongbin Li (Tongyi Lab, Alibaba) introduce the Causal Improvement Graph (CIG), which stores the state of an automated harness-optimization search in a persistent graph instead of in the proposer's growing history.

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
Mohamed Dhouib, Sonia Vanier and colleagues at École polytechnique (LIX), with Elie Bursztein of Google DeepMind, introduce RAISED, a self-distillation defense that trains a tool-using agent to behave on injected contexts exactly as it behaves on clean ones.

AECP: Artifact-Exclusive Communication Protocol for Multi-Agent Code Generation
Jiaqi Xue (UCF, during an AWS internship) with Yanjun Wang, Myeongsoo Kim and colleagues at AWS AI Labs propose AECP, a protocol in which coding agents communicate only through structured artifacts that the harness processes.

Harness-Search: Guiding Long-Horizon Search through Multi-Agent Coordination
Shanyong Wang, Yanyu Xu and colleagues at Xiaohongshu with Shandong University and UIUC introduce Harness-Search, which splits a long-horizon search agent into three permission-bounded roles that propose, commit and audit.

Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks
Reza Esfandiarpoor, Radek Osmulski, Even Oldridge and colleagues at NVIDIA (with the University of Edinburgh) measure what a ReAct retrieval agent adds over dense retrieval on complex search, and what it costs.