🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
Coding-Agent Benchmarks Should Match Their Users' Task Flows

Coding-Agent Benchmarks Should Match Their Users' Task Flows

Igor Slinko, Yaroslav Golubev and Sergey Titov (JetBrains Research) compare 4,782 real coding-agent sessions from JetBrains IDEs with issue-derived benchmarks and propose SWE-TaskFlow to reshape benchmarks toward a measured interaction pattern.

25Evaluation
Agent Plasticity: Measuring Self-Improvement Through Experience

Agent Plasticity: Measuring Self-Improvement Through Experience

Harman Singh, Anirudh Goyal, Jason Weston, Gabriel Synnaeve, Rob Fergus and colleagues at Meta Superintelligence Labs (with UC Berkeley and Princeton's Sanjeev Arora) propose agent plasticity, a measure of how efficiently an agent turns experience into gains on held-out tasks.

26Agents
Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver

Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver

Egor Pakhomov and Erik Nijkamp (Salesforce AI Research) test what MemoryAgentBench's Conflict Resolution split measures by running its own stated rule, newest statement wins, as a zero-learning resolver. Accepted at the NeurIPS 2026 Interpreting Agent Behavior workshop.

27Agents
Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

Panagiotis Kasnesis and colleagues (University of West Attica) evaluate 9 models from 0.8B parameters to a hosted frontier model at each of the five LLM call sites of Wactorz, a deployed open-source multi-agent home-automation framework. Accepted at a NeurIPS 2026 workshop on SLMs for agentic systems.

28Agents
SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

Yuyao Ge and colleagues at the Institute of Computing Technology, Chinese Academy of Sciences (with UC Merced and Tsinghua) present SkillForge, an agentic RL method in which the skill library and the policy are updated together. Accepted at NeurIPS 2026.

29Agents
DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents

DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents

Jike Zhong, Ritwick Chaudhry, Nishant Sankaran and colleagues at Amazon AGI (with USC) introduce DSV-Mem, a benchmark for dense, stateful visual memory in multimodal agents that assist with professional workflows.

30Memory
From Evidence to Action: How Tool-Using Agents Fail

From Evidence to Action: How Tool-Using Agents Fail

Hongzhan Lin, Shidong Cao, Ziyang Luo and Wenhao Chai (Princeton) with Mong-Li Lee and Wynne Hsu (NUS) and authors at HKBU and Amazon Web Services study where tool-using agents break the chain from established evidence to action, and release SafeActBench.

31Agents
ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?

ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?

Haizhong Zheng, Yizhuo Di, Ranajoy Sadhukhan, Shuowei Jin and Beidi Chen (Carnegie Mellon, Infini-AI Lab) introduce ServeLearnBench, a benchmark for agents that must infer and revise hidden environment policies from serving experience.

32Agents
Structuring MoE Expert Selection for Agentic Reinforcement Learning

Structuring MoE Expert Selection for Agentic Reinforcement Learning

Bolian Li (Apple and Purdue) with Ting-Yao Hu, Cheng-Yu Hsieh, Oncel Tuzel, Raviteja Vemulapalli and colleagues at Apple study how mixture-of-experts routing relates to agentic behavior and control it during RL post-training.

33Architecture
Stateless Language Agents: Scaling Long-Horizon Automated Research

Stateless Language Agents: Scaling Long-Horizon Automated Research

Qizheng Zhang, Changxiu Ji, Kunle Olukotun and colleagues at Stanford, with CMU, UW and SambaNova, introduce Stateless Language Agents (SLA), where the harness holds all research state and every agent call starts from a freshly built context.

34Agents
ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

Yupeng Su, Jiayi Tian and Zheng Zhang (UC Santa Barbara) with Souvik Kundu (Intel) present ReFold, a training-free rendering layer that compresses what a long-horizon agent sees while keeping the full interaction history recoverable.

35Agents
When to Remember, When to Abstain: Category-Conditioned Retention for Reliable Agent Memory

When to Remember, When to Abstain: Category-Conditioned Retention for Reliable Agent Memory

Olukunle Owolabi, Pulkit Gupta and Fei Wang at Meta AI study when an agent's memory pipeline should store an inferred assertion, and propose confidence thresholds conditioned on the assertion's semantic category. The paper is accepted at the NeurIPS 2026 Social Agent Workshop.

36Agents
AMBER: Training Long-Horizon Web Agents through Append-Only Memory

AMBER: Training Long-Horizon Web Agents through Append-Only Memory

Chinmay Savadikar, Tianfu Wu, Lingyun Wang and colleagues at North Carolina State University and Shopify introduce AMBER, an append-only memory that a web agent learns to write while it reasons and acts, trained end-to-end with outcome-reward RL.

37Memory
NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale

NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale

Songlin Jiang and Mario Di Francesco (Aalto University) with Zhiyu Li, Terry Kong and colleagues at NVIDIA present NeMo-DCR, a bit-exact delta-compressed weight synchronization (refit) for agentic RL, released in NVIDIA NeMo RL.

38Agents
Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale

Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale

Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro and Lorenzo Dini at Google describe FlowAgent, a program repair agent deployed across Google that fixes test failures during pre-submit continuous integration, before the developer moves on to other work. The paper is accepted at ASE 2026.

39Agents
SquidAgent: Parallelize Wisely, Coordinate Efficiently

SquidAgent: Parallelize Wisely, Coordinate Efficiently

Yexiong Lin, Tongliang Liu and colleagues at the University of Sydney with Shanshan Ye (MBZUAI) present SquidAgent, which decides layer by layer whether to parallelize agent work; the paper is accepted at NeurIPS 2026.

40Agents
Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations

Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations

Toby D. Pilditch, Konstantinos Voudouris, Alexandra Abbas and Cozmin Ududec at the UK AI Security Institute release Transect, an open-source package built on Inspect Scout for analysing very long agent evaluation transcripts in a reproducible way.

41Agents
MemTrace: State-Consistent Memory for Long-Horizon Coding Agents

MemTrace: State-Consistent Memory for Long-Horizon Coding Agents

Hongming Xu, Zhiyu Li, Juncheng Zhang and colleagues at Shanghai Jiao Tong University and MemTensor introduce MemTrace, a memory system for long-horizon coding agents that checks recalled evidence against the current repository state before reusing it.

42Code
MLLMs Fail to Refuse when Using Tools Agentically

MLLMs Fail to Refuse when Using Tools Agentically

Rikiya Takehi (MIT, during an internship at NVIDIA) with Ryo Hachiuma, Shaona Ghosh and colleagues at NVIDIA show that giving multimodal LLMs tools makes them refuse harmful requests less often; the paper is accepted at NeurIPS 2026.

43Agents
Causal Improvement Graph for Agentic Harness Optimization

Causal Improvement Graph for Agentic Harness Optimization

Junjie Zhang, Shunyu Liu and Dacheng Tao (Nanyang Technological University) with Ting-En Lin and Yongbin Li (Tongyi Lab, Alibaba) introduce the Causal Improvement Graph (CIG), which stores the state of an automated harness-optimization search in a persistent graph instead of in the proposer's growing history.

44Agents
RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

Mohamed Dhouib, Sonia Vanier and colleagues at École polytechnique (LIX), with Elie Bursztein of Google DeepMind, introduce RAISED, a self-distillation defense that trains a tool-using agent to behave on injected contexts exactly as it behaves on clean ones.

45Agents
AECP: Artifact-Exclusive Communication Protocol for Multi-Agent Code Generation

AECP: Artifact-Exclusive Communication Protocol for Multi-Agent Code Generation

Jiaqi Xue (UCF, during an AWS internship) with Yanjun Wang, Myeongsoo Kim and colleagues at AWS AI Labs propose AECP, a protocol in which coding agents communicate only through structured artifacts that the harness processes.

46Agents
Harness-Search: Guiding Long-Horizon Search through Multi-Agent Coordination

Harness-Search: Guiding Long-Horizon Search through Multi-Agent Coordination

Shanyong Wang, Yanyu Xu and colleagues at Xiaohongshu with Shandong University and UIUC introduce Harness-Search, which splits a long-horizon search agent into three permission-bounded roles that propose, commit and audit.

47Agents
Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks

Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks

Reza Esfandiarpoor, Radek Osmulski, Even Oldridge and colleagues at NVIDIA (with the University of Edinburgh) measure what a ReAct retrieval agent adds over dense retrieval on complex search, and what it costs.

48Retrieval
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026