🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning

PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning

Jiaan Zhu, Wei Gao and colleagues at USTC and HKUST with Alibaba Group present PEARL, an asynchronous agentic RL system that combines elastic GPUs, temporary reuse of idle training GPUs, and per-workload choice between prefill-decode colocation and disaggregation to speed up multi-turn rollouts.

97Agents
ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning

ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning

Kun Feng, Yuchen Fang and colleagues at ShanghaiTech University and Ant Group introduce ARISE, an agentic RL framework that turns rollout evidence into paired rubrics and skills, retires criteria once mastered, and samples tasks by estimated capability, raising Qwen3.5-27B from 23.4% to 45.6% on SkillsBench.

98Agents
"You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely Accused

"You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely Accused

Xutao Mao, Rui Qian and colleagues at City University of Hong Kong and Fudan introduce CAVE-Bench, 365 agentic tasks that test whether an agent keeps verified correct work when a later message falsely blames it for a failure, and find that agents damage correct work in up to 60% of runs.

99Agents
Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents

Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents

Ritul Satish, Prasoon Sinha, Akiho Kawada and Neeraja Yadwadkar at UT Austin run nearly 35,000 agent runs on SWE-bench Verified and Terminal-Bench to separate the three decisions in a context compression policy (mechanism, trigger, and amount removed) and measure how each affects success, tokens, latency and cost.

100Efficiency
EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks

EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks

Mukul Singh and colleagues at Microsoft introduce EmailBench, a self-contained benchmark of 206 enterprise email and productivity scenarios built on a typed email API and a synthetic Enron-style corpus, and show that agents complete most of their tool calls correctly while failing most tasks.

101Evaluation
Just-In-Time Agent Memory with Runtime Agentic Research

Just-In-Time Agent Memory with Runtime Agentic Research

Bingyu Yan, Zheng Liu and colleagues at the Beijing Academy of Artificial Intelligence propose Just-In-Time Agent Memory (JAM), which keeps complete raw histories and builds query-specific context at request time with a trained Researcher agent, instead of compressing memory before requests arrive.

102Agents
FlowState: Execution State as Memory for Long-Horizon LLM Agents

FlowState: Execution State as Memory for Long-Horizon LLM Agents

Minghao Li, Bangyan Li and colleagues at Ant International, Ant Group propose FlowState, an agent memory framework that stores execution state as typed, linked nodes with references back to raw tool output, so an agent can revisit earlier decisions and their evidence without carrying the full history in context.

103Agents
CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering

CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering

Prince Zizhuang Wang (CMU) and colleagues at USC, UW-Madison and other universities introduce CUA-SWE, a benchmark and environment where agents must both edit code and operate the running application through its GUI to diagnose, fix and verify bugs across web, game, mobile and DevOps projects.

104Agents
Inspire: Benchmarking Scientific Literature Search for Open Research Problems

Inspire: Benchmarking Scientific Literature Search for Open Research Problems

Jianrong Ding (CUHK, Microsoft Research Asia intern) and colleagues at Microsoft Research Asia and CUHK introduce INSPIRE, a benchmark where agents search an open, date-gated corpus for prior work that later solved a redacted research problem, scored separately on exposure, selection and ranking.

105Evaluation
LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles

LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles

Bingo Zhang and colleagues at Vera Praxis with Tencent, HKUST and CUHK introduce LongPuzzleBench, 114 levels of six visual puzzle games played only through GUI input, where a legal move can make a level unsolvable without any signal until several moves later.

106Agents
Agents as Software: A Programming Languages Agenda for Agent Reliability

Agents as Software: A Programming Languages Agenda for Agent Reliability

Shraddha Barke (Microsoft Research) and Adithya Murali (UW-Madison) argue in an Onward! 2026 essay that agents should be treated as programs whose behavior is spread across prompts, tools, memories and traces, and lay out how specifications, static analysis and runtime monitoring from programming languages research apply to them.

107Agents
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Fuli Luo and colleagues at Xiaomi's LLM Core team, with Renmin, Peking and HKU, introduce GAGAR, which uses an agentic grader to rank test-passing trajectories within each RL rollout group and redistributes advantage toward cleaner, more targeted patches.

108Agents
When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents

When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents

Janvijay Singh (UIUC, Microsoft Research intern), Vaishnavi Shrivastava, Dilek Hakkani-Tür, Ece Kamar and Asli Celikyilmaz at Microsoft Research AI Frontiers introduce AGNI, a pipeline that breaks one environmental assumption behind a successful terminal-agent trajectory while keeping the task solvable, and use it to measure how well agents adapt.

109Agents
Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

Michael Hardy, Anka Reuel, Mykel Kochenderfer and Sanmi Koyejo at Stanford (with UIUC) build a Bayesian variance-decomposition framework for sparse agent leaderboards and apply it to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index.

110Evaluation
Self-Evolving Coding Rules for AI Coding Agents

Self-Evolving Coding Rules for AI Coding Agents

Zhengyuan Jiang, Neil Zhenqiang Gong and colleagues at Duke University introduce RuleEvolve (NeurIPS 2026), which evolves the coding-rules files that coding agents read instead of relying on hand-written ones.

111Code
Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

Sahan Paliskara (independent), Nattaput Namchittai (Stanford), Andrew Lampinen (Anthropic) and colleagues study what happens when several agents, each acting for a different user, share a resource such as a compute budget, a calendar or a release cutoff.

112Agents
Finding the Right Fit: Model-Harness Interactions across Agent Tasks

Finding the Right Fit: Model-Harness Interactions across Agent Tasks

Yixuan Li, Bo An and colleagues at Nanyang Technological University evaluate 66 model-harness configurations and show that model rankings, best harnesses and cost-efficiency all change with the harness and the benchmark, so the pairing has to be evaluated as a unit.

113Evaluation
ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Sungho Park (POSTECH, intern at Microsoft), Jue Zhang, Pengfei Gao and colleagues at Microsoft, POSTECH and KAIST introduce ActiveSaddler, which adapts the training scenarios used to drive automated harness optimization as the harness changes.

114Agents
DAYJOB: A Benchmark for Long-Horizon Professional Work

DAYJOB: A Benchmark for Long-Horizon Professional Work

Stephanie Finley, Liudas Panavas, Sushant Mehta, Edwin Chen and colleagues at Surge AI release DAYJOB, 130 healthcare and finance tasks written by working professionals, each estimated at 13 to 17 hours of human work and graded all-or-nothing against an expert rubric.

115Evaluation
DeFA: Dependency-Guided Failure Attribution for LLM Agents

DeFA: Dependency-Guided Failure Attribution for LLM Agents

Bo Deng, Kang Zhou, Lifan Guo and colleagues at Qwen DianJin Team (Alibaba Cloud) with Beihang University introduce DeFA, which builds a dependency graph over an agent trajectory and traces how errors propagate to find the decisive error, the responsible agent and its category.

116Agents
My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning

My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning

Yihua Zhu, Qianying Liu, Weixu Qiao and colleagues at Alibaba with Kyoto University propose FAULT, which converts an agent's own natural-language diagnosis of its errors into step-level credit that is anchored to the terminal reward in agentic RL.

117Agents
AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

AutoCompact trains a coding agent to decide when to compact its context, what working state to keep, and how to continue afterward, as part of its own policy. A judge reviews the base agent's compaction decisions and replaces flawed ones before they execute, and the corrected trajectories are used for SFT and then for RL that optimizes coding and compaction together on task success. Pass rates rise by 9.2 points on SWE-bench Verified and 5.0 points on SWE-PolyBench Verified, and the gains hold both with a 256K window that never overflows and with a 16K window that falls back to forced compaction.

118Code
Harness Learning Enables Generalizable Test-Time Adaptation

Harness Learning Enables Generalizable Test-Time Adaptation

Alvin Zhang, Xuecheng Liu, Zixuan Wang, Ruslan Salakhutdinov, Daniel Khashabi, Yuda Song, Andrea Zanette and colleagues at Carnegie Mellon University train a proposer model with RL to edit an agent's executable harness from execution feedback, and show the learned revision skill transfers to tasks it never saw.

119Agents
GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

Geyi Yang, Zhongxiang Dai and colleagues at CUHK-Shenzhen, Tianjin University, HIT Shenzhen and ECNU present GUI-HARVEST, an automatic harness optimizer that improves GUI agents with frozen backbones by grounding failure diagnosis in screenshots and repeated runs.

120Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026