AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Coding Agents are Strong Prompt Optimizers
Search-based prompt optimizers such as GEPA propose edits, run fresh rollouts, score them, and keep only the edits that improve a validation metric. Researchers from Microsoft show that this loop may be unnecessary when you already have a corpus of agent trajectories.

Agensh: Scaling Organizational Intelligence to 1,024 Agents
Multi-agent harnesses usually depend on a central orchestrator that assigns tasks and coordinates workers, and that orchestrator limits how many agents the system can use. Microsoft Research introduces Agensh, a self-organized multi-agent harness with no central orchestrator, and scales it to 1,024 coding agents.

Recursive self-improvement of AI research agents
Dhruv Srikanth, Zhengyao Jiang and colleagues at Weco AI present AIDE^2, a loop in which a frontier AI research agent edits its own code, benchmarks the new versions on AI R&D tasks and keeps the edits that score best on hidden evaluations.

Emergent Collusion in Long-Horizon LLM Agent Interaction
Xinrui Shi and Diyi Yang (Stanford) with Yanzhe Zhang (Georgia Tech) show that two LLM agents that repeatedly verify each other's work drift into collusion when following the verification protocol conflicts with maximizing reward.

VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks
Liyang Fan, Min Yang, Jieping Ye and colleagues at SIAT (Chinese Academy of Sciences), SUAT and Alibaba build VibeMemBench to measure whether memory systems improve coding agents on executable repository tasks, and find that current systems mostly do not.

MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents
Ruike Cao, Fanyu Zhao and colleagues at Alibaba's Qwen Applications Business Group, USTC and Fudan introduce MemCalib, a benchmark for whether a model gives each retrieved memory the right amount of influence on its answer, and MemCalib-RL to train that behavior.

Toollery: Scaling LLM Agents to Thousands of Skills and Tools
Xiangxi Tian and Ran Guan (Huawei 2012 Laboratories) present Toollery, a training-free way to narrow libraries of thousands of skills and tools to a short candidate list before the LLM makes its final selection.

SelfOp: An Optimization Algorithm for Self-Improving Security Agents
Saad Ullah and Gianluca Stringhini (Boston University) with Yigitcan Kaya, Christopher Kruegel and Giovanni Vigna (UC Santa Barbara) present SelfOp, which improves a frozen security agent by editing its instructions, skills and reference documents through textual gradient descent.

One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents
Jie Zhao, Ziyu Jiang and colleagues at Alibaba Group (Logics team) find that pooled agentic RL on SWE tasks improves some task categories while regressing others, and train per-category experts that are merged back into one Qwen3.6-27B policy.

A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents
Aakash Kolekar, Sahika Genc and colleagues at Amazon Advertising and AWS Agentic AI study when to use SFT, RL or both for long-horizon advertising analytics agents, and turn the answer into a per-feature routing diagnostic (EMNLP 2026 Industry Track).

FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model
Jingxuan Xu, Gang Wu, Yanan Wu, Yutao Mou and colleagues (independent researchers with Peking, Nanjing and BUPT) propose FLARE, which trains a lightweight generative reward model to give step-level risk feedback to long-horizon coding agents during inference, SFT and RL.

Harness-Zero: Harness Distillation via Agent-as-Harness
A specialized harness can raise an agent's performance a lot, but the best harness differs across domains, instances, and models. Harness-Zero, from Google and colleagues, uses the specialized harness only during training and moves the behavior it induces into the model weights.

XYEval: Agents say yes to bad advice
Users often suggest a fix that sounds right and is wrong, and Google DeepMind's XYEval measures how often agents go along with it by adding one confident, misleading hint to tasks from tau2-bench, SWE-bench, Terminal-Bench, HLE, and MCP-Atlas while keeping the correct solution unchanged. Scores fall by up to 46.7% relative across Gemini, Claude Opus 4.8, and GPT 5.5, and agents often disagree with the hint in their reasoning and then follow it without telling the user. A system prompt warning about the XY problem helps on single-turn tasks but leaves large drops on multi-turn ones such as tau2-bench and SWE-bench Verified.

Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents
Many agentic memory systems use an autoregressive LLM to decide how memories are organized, retrieved, and used, which puts expensive generation on the critical path of every memory operation. Jev-Mem borrows its design from System-One/System-Two cognition and hands those decisions to a lightweight controller.

Beyond Task Completion: Training Capable and Safe Computer-Use Agents
Zeyu Kang, Xinquan Chen, Xuhong Wang and colleagues at Shanghai AI Laboratory train a computer-use agent to finish benign tasks, work around hazards when a safe path exists, and refuse harmful goals, using one joint SFT-then-RL recipe called SCOPE.

An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents
Luzhuo Chen and Jiayu Shi (Paritok) instrument their production compression gateway between Claude Code or Codex and Claude Sonnet or GPT-5, and separate the token bill into three levers that save at very different rates. The authors build the gateway being measured.

When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain
Yunxiang Li, Xixin Wu and Helen Meng (CUHK) propose CIGAsk, an RL recipe that teaches a model both when to ask a clarifying question and how to phrase one that recovers the missing information (EMNLP 2026 Findings).

Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents
Kaijie Chen, Chenyu Fang, Peng Ye and colleagues at Tongji University, Shanghai AI Laboratory and Fudan propose Trace, which compiles noisy sparse-reward trajectories into short, state-conditioned procedures that an agent can execute and verify.

CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents
Xinlu Zhang, Besnik Fetahu, Xi Chen and colleagues at Amazon show that a search agent trained with GRPO under one harness learns parallel search only for that harness, and propose CHART, a rotating harness curriculum that makes the behavior hold across prompt rewrites (NeurIPS 2026 CL4FMAgents workshop).

PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents
Pirzada Suhail, Menglin Xia and colleagues at Microsoft Research and M365 propose Pseudo Self-Distillation (PSD), which trains small Qwen3 models to run a multi-stage memory-construction pipeline that normally needs GPT-4.1-mini, using only the oracle's text outputs.

DolphinBench: Mapping the Pareto Frontier of Agent Memory
Soumil Rathi, Deshraj Yadav and Taranjeet Singh (Mem0) release DolphinBench, a memory benchmark that scores agents on actions they take in simulated apps rather than on answers to recall questions, and requires every submission to report cost and latency with accuracy.

Learning Generalizable Behaviors for Terminal Agents
Yihang Yao, Bo Pang, Semih Yavuz and colleagues at Salesforce AI Research, with Ding Zhao at Carnegie Mellon, study what RL actually changes in terminal agents and propose RIVER, a training recipe that improves reward quality by filtering defective environments and penalizing repetitive loops.

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
Peng Xia, Chen-Yu Lee, Tomas Pfister and colleagues at Google Cloud AI Research, with UNC Chapel Hill, Stanford and WashU, show that automated harness evolution overfits its training tasks and add regularization to both the edit proposer and the selector.

Self-Organizing Agent Teams Learn to Reason Together
Multi-agent systems usually fix roles and protocols in advance. Researchers from Stanford and Together AI let a fixed team of models learn how to organize its own collaboration from past exchanges.