AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning
Jiaan Zhu, Wei Gao and colleagues at USTC and HKUST with Alibaba Group present PEARL, an asynchronous agentic RL system that combines elastic GPUs, temporary reuse of idle training GPUs, and per-workload choice between prefill-decode colocation and disaggregation to speed up multi-turn rollouts.

ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning
Kun Feng, Yuchen Fang and colleagues at ShanghaiTech University and Ant Group introduce ARISE, an agentic RL framework that turns rollout evidence into paired rubrics and skills, retires criteria once mastered, and samples tasks by estimated capability, raising Qwen3.5-27B from 23.4% to 45.6% on SkillsBench.

"You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely Accused
Xutao Mao, Rui Qian and colleagues at City University of Hong Kong and Fudan introduce CAVE-Bench, 365 agentic tasks that test whether an agent keeps verified correct work when a later message falsely blames it for a failure, and find that agents damage correct work in up to 60% of runs.

Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents
Ritul Satish, Prasoon Sinha, Akiho Kawada and Neeraja Yadwadkar at UT Austin run nearly 35,000 agent runs on SWE-bench Verified and Terminal-Bench to separate the three decisions in a context compression policy (mechanism, trigger, and amount removed) and measure how each affects success, tokens, latency and cost.

EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks
Mukul Singh and colleagues at Microsoft introduce EmailBench, a self-contained benchmark of 206 enterprise email and productivity scenarios built on a typed email API and a synthetic Enron-style corpus, and show that agents complete most of their tool calls correctly while failing most tasks.

Just-In-Time Agent Memory with Runtime Agentic Research
Bingyu Yan, Zheng Liu and colleagues at the Beijing Academy of Artificial Intelligence propose Just-In-Time Agent Memory (JAM), which keeps complete raw histories and builds query-specific context at request time with a trained Researcher agent, instead of compressing memory before requests arrive.

FlowState: Execution State as Memory for Long-Horizon LLM Agents
Minghao Li, Bangyan Li and colleagues at Ant International, Ant Group propose FlowState, an agent memory framework that stores execution state as typed, linked nodes with references back to raw tool output, so an agent can revisit earlier decisions and their evidence without carrying the full history in context.

CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering
Prince Zizhuang Wang (CMU) and colleagues at USC, UW-Madison and other universities introduce CUA-SWE, a benchmark and environment where agents must both edit code and operate the running application through its GUI to diagnose, fix and verify bugs across web, game, mobile and DevOps projects.

Inspire: Benchmarking Scientific Literature Search for Open Research Problems
Jianrong Ding (CUHK, Microsoft Research Asia intern) and colleagues at Microsoft Research Asia and CUHK introduce INSPIRE, a benchmark where agents search an open, date-gated corpus for prior work that later solved a redacted research problem, scored separately on exposure, selection and ranking.

LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles
Bingo Zhang and colleagues at Vera Praxis with Tencent, HKUST and CUHK introduce LongPuzzleBench, 114 levels of six visual puzzle games played only through GUI input, where a legal move can make a level unsolvable without any signal until several moves later.

Agents as Software: A Programming Languages Agenda for Agent Reliability
Shraddha Barke (Microsoft Research) and Adithya Murali (UW-Madison) argue in an Onward! 2026 essay that agents should be treated as programs whose behavior is spread across prompts, tools, memories and traces, and lay out how specifications, static analysis and runtime monitoring from programming languages research apply to them.

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Fuli Luo and colleagues at Xiaomi's LLM Core team, with Renmin, Peking and HKU, introduce GAGAR, which uses an agentic grader to rank test-passing trajectories within each RL rollout group and redistributes advantage toward cleaner, more targeted patches.

When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents
Janvijay Singh (UIUC, Microsoft Research intern), Vaishnavi Shrivastava, Dilek Hakkani-Tür, Ece Kamar and Asli Celikyilmaz at Microsoft Research AI Frontiers introduce AGNI, a pipeline that breaks one environmental assumption behind a successful terminal-agent trajectory while keeping the task solvable, and use it to measure how well agents adapt.

Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard
Michael Hardy, Anka Reuel, Mykel Kochenderfer and Sanmi Koyejo at Stanford (with UIUC) build a Bayesian variance-decomposition framework for sparse agent leaderboards and apply it to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index.

Self-Evolving Coding Rules for AI Coding Agents
Zhengyuan Jiang, Neil Zhenqiang Gong and colleagues at Duke University introduce RuleEvolve (NeurIPS 2026), which evolves the coding-rules files that coding agents read instead of relying on hand-written ones.

Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams
Sahan Paliskara (independent), Nattaput Namchittai (Stanford), Andrew Lampinen (Anthropic) and colleagues study what happens when several agents, each acting for a different user, share a resource such as a compute budget, a calendar or a release cutoff.

Finding the Right Fit: Model-Harness Interactions across Agent Tasks
Yixuan Li, Bo An and colleagues at Nanyang Technological University evaluate 66 model-harness configurations and show that model rankings, best harnesses and cost-efficiency all change with the harness and the benchmark, so the pairing has to be evaluated as a unit.

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
Sungho Park (POSTECH, intern at Microsoft), Jue Zhang, Pengfei Gao and colleagues at Microsoft, POSTECH and KAIST introduce ActiveSaddler, which adapts the training scenarios used to drive automated harness optimization as the harness changes.

DAYJOB: A Benchmark for Long-Horizon Professional Work
Stephanie Finley, Liudas Panavas, Sushant Mehta, Edwin Chen and colleagues at Surge AI release DAYJOB, 130 healthcare and finance tasks written by working professionals, each estimated at 13 to 17 hours of human work and graded all-or-nothing against an expert rubric.

DeFA: Dependency-Guided Failure Attribution for LLM Agents
Bo Deng, Kang Zhou, Lifan Guo and colleagues at Qwen DianJin Team (Alibaba Cloud) with Beihang University introduce DeFA, which builds a dependency graph over an agent trajectory and traces how errors propagate to find the decisive error, the responsible agent and its category.

My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning
Yihua Zhu, Qianying Liu, Weixu Qiao and colleagues at Alibaba with Kyoto University propose FAULT, which converts an agent's own natural-language diagnosis of its errors into step-level credit that is anchored to the terminal reward in agentic RL.

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
AutoCompact trains a coding agent to decide when to compact its context, what working state to keep, and how to continue afterward, as part of its own policy. A judge reviews the base agent's compaction decisions and replaces flawed ones before they execute, and the corrected trajectories are used for SFT and then for RL that optimizes coding and compaction together on task success. Pass rates rise by 9.2 points on SWE-bench Verified and 5.0 points on SWE-PolyBench Verified, and the gains hold both with a 256K window that never overflows and with a 16K window that falls back to forced compaction.

Harness Learning Enables Generalizable Test-Time Adaptation
Alvin Zhang, Xuecheng Liu, Zixuan Wang, Ruslan Salakhutdinov, Daniel Khashabi, Yuda Song, Andrea Zanette and colleagues at Carnegie Mellon University train a proposer model with RL to edit an agent's executable harness from execution feedback, and show the learned revision skill transfers to tasks it never saw.

GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution
Geyi Yang, Zhongxiang Dai and colleagues at CUHK-Shenzhen, Tianjin University, HIT Shenzhen and ECNU present GUI-HARVEST, an automatic harness optimizer that improves GUI agents with frozen backbones by grounding failure diagnosis in screenshots and repeated runs.