🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
Coding Agents are Strong Prompt Optimizers

Coding Agents are Strong Prompt Optimizers

Search-based prompt optimizers such as GEPA propose edits, run fresh rollouts, score them, and keep only the edits that improve a validation metric. Researchers from Microsoft show that this loop may be unnecessary when you already have a corpus of agent trajectories.

241Agents
Agensh: Scaling Organizational Intelligence to 1,024 Agents

Agensh: Scaling Organizational Intelligence to 1,024 Agents

Multi-agent harnesses usually depend on a central orchestrator that assigns tasks and coordinates workers, and that orchestrator limits how many agents the system can use. Microsoft Research introduces Agensh, a self-organized multi-agent harness with no central orchestrator, and scales it to 1,024 coding agents.

242Agents
Recursive self-improvement of AI research agents

Recursive self-improvement of AI research agents

Dhruv Srikanth, Zhengyao Jiang and colleagues at Weco AI present AIDE^2, a loop in which a frontier AI research agent edits its own code, benchmarks the new versions on AI R&D tasks and keeps the edits that score best on hidden evaluations.

243Agents
Emergent Collusion in Long-Horizon LLM Agent Interaction

Emergent Collusion in Long-Horizon LLM Agent Interaction

Xinrui Shi and Diyi Yang (Stanford) with Yanzhe Zhang (Georgia Tech) show that two LLM agents that repeatedly verify each other's work drift into collusion when following the verification protocol conflicts with maximizing reward.

244Agents
VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

Liyang Fan, Min Yang, Jieping Ye and colleagues at SIAT (Chinese Academy of Sciences), SUAT and Alibaba build VibeMemBench to measure whether memory systems improve coding agents on executable repository tasks, and find that current systems mostly do not.

245Code
MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

Ruike Cao, Fanyu Zhao and colleagues at Alibaba's Qwen Applications Business Group, USTC and Fudan introduce MemCalib, a benchmark for whether a model gives each retrieved memory the right amount of influence on its answer, and MemCalib-RL to train that behavior.

246Evaluation
Toollery: Scaling LLM Agents to Thousands of Skills and Tools

Toollery: Scaling LLM Agents to Thousands of Skills and Tools

Xiangxi Tian and Ran Guan (Huawei 2012 Laboratories) present Toollery, a training-free way to narrow libraries of thousands of skills and tools to a short candidate list before the LLM makes its final selection.

247Agents
SelfOp: An Optimization Algorithm for Self-Improving Security Agents

SelfOp: An Optimization Algorithm for Self-Improving Security Agents

Saad Ullah and Gianluca Stringhini (Boston University) with Yigitcan Kaya, Christopher Kruegel and Giovanni Vigna (UC Santa Barbara) present SelfOp, which improves a frozen security agent by editing its instructions, skills and reference documents through textual gradient descent.

248Agents
One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents

One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents

Jie Zhao, Ziyu Jiang and colleagues at Alibaba Group (Logics team) find that pooled agentic RL on SWE tasks improves some task categories while regressing others, and train per-category experts that are merged back into one Qwen3.6-27B policy.

249Code
A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents

A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents

Aakash Kolekar, Sahika Genc and colleagues at Amazon Advertising and AWS Agentic AI study when to use SFT, RL or both for long-horizon advertising analytics agents, and turn the answer into a per-feature routing diagnostic (EMNLP 2026 Industry Track).

250Agents
FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model

FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model

Jingxuan Xu, Gang Wu, Yanan Wu, Yutao Mou and colleagues (independent researchers with Peking, Nanjing and BUPT) propose FLARE, which trains a lightweight generative reward model to give step-level risk feedback to long-horizon coding agents during inference, SFT and RL.

251Reinforcement Learning
Harness-Zero: Harness Distillation via Agent-as-Harness

Harness-Zero: Harness Distillation via Agent-as-Harness

A specialized harness can raise an agent's performance a lot, but the best harness differs across domains, instances, and models. Harness-Zero, from Google and colleagues, uses the specialized harness only during training and moves the behavior it induces into the model weights.

252Agents
XYEval: Agents say yes to bad advice

XYEval: Agents say yes to bad advice

Users often suggest a fix that sounds right and is wrong, and Google DeepMind's XYEval measures how often agents go along with it by adding one confident, misleading hint to tasks from tau2-bench, SWE-bench, Terminal-Bench, HLE, and MCP-Atlas while keeping the correct solution unchanged. Scores fall by up to 46.7% relative across Gemini, Claude Opus 4.8, and GPT 5.5, and agents often disagree with the hint in their reasoning and then follow it without telling the user. A system prompt warning about the XY problem helps on single-turn tasks but leaves large drops on multi-turn ones such as tau2-bench and SWE-bench Verified.

253Agents
Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Many agentic memory systems use an autoregressive LLM to decide how memories are organized, retrieved, and used, which puts expensive generation on the critical path of every memory operation. Jev-Mem borrows its design from System-One/System-Two cognition and hands those decisions to a lightweight controller.

254Agents
Beyond Task Completion: Training Capable and Safe Computer-Use Agents

Beyond Task Completion: Training Capable and Safe Computer-Use Agents

Zeyu Kang, Xinquan Chen, Xuhong Wang and colleagues at Shanghai AI Laboratory train a computer-use agent to finish benign tasks, work around hazards when a safe path exists, and refuse harmful goals, using one joint SFT-then-RL recipe called SCOPE.

255Agents
An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents

An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents

Luzhuo Chen and Jiayu Shi (Paritok) instrument their production compression gateway between Claude Code or Codex and Claude Sonnet or GPT-5, and separate the token bill into three levers that save at very different rates. The authors build the gateway being measured.

256Efficiency
When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain

When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain

Yunxiang Li, Xixin Wu and Helen Meng (CUHK) propose CIGAsk, an RL recipe that teaches a model both when to ask a clarifying question and how to phrase one that recovers the missing information (EMNLP 2026 Findings).

257Agents
Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents

Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents

Kaijie Chen, Chenyu Fang, Peng Ye and colleagues at Tongji University, Shanghai AI Laboratory and Fudan propose Trace, which compiles noisy sparse-reward trajectories into short, state-conditioned procedures that an agent can execute and verify.

258Agents
CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents

CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents

Xinlu Zhang, Besnik Fetahu, Xi Chen and colleagues at Amazon show that a search agent trained with GRPO under one harness learns parallel search only for that harness, and propose CHART, a rotating harness curriculum that makes the behavior hold across prompt rewrites (NeurIPS 2026 CL4FMAgents workshop).

259Agents
PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents

PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents

Pirzada Suhail, Menglin Xia and colleagues at Microsoft Research and M365 propose Pseudo Self-Distillation (PSD), which trains small Qwen3 models to run a multi-stage memory-construction pipeline that normally needs GPT-4.1-mini, using only the oracle's text outputs.

260Memory
DolphinBench: Mapping the Pareto Frontier of Agent Memory

DolphinBench: Mapping the Pareto Frontier of Agent Memory

Soumil Rathi, Deshraj Yadav and Taranjeet Singh (Mem0) release DolphinBench, a memory benchmark that scores agents on actions they take in simulated apps rather than on answers to recall questions, and requires every submission to report cost and latency with accuracy.

261Agents
Learning Generalizable Behaviors for Terminal Agents

Learning Generalizable Behaviors for Terminal Agents

Yihang Yao, Bo Pang, Semih Yavuz and colleagues at Salesforce AI Research, with Ding Zhao at Carnegie Mellon, study what RL actually changes in terminal agents and propose RIVER, a training recipe that improves reward quality by filtering defective environments and penalizing repetitive loops.

262Agents
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Peng Xia, Chen-Yu Lee, Tomas Pfister and colleagues at Google Cloud AI Research, with UNC Chapel Hill, Stanford and WashU, show that automated harness evolution overfits its training tasks and add regularization to both the edit proposer and the selector.

263Agents
Self-Organizing Agent Teams Learn to Reason Together

Self-Organizing Agent Teams Learn to Reason Together

Multi-agent systems usually fix roles and protocols in advance. Researchers from Stanford and Together AI let a fixed team of models learn how to organize its own collaboration from past exchanges.

264Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026