🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

Yifan Zhang, Yutong Dai, Ran Xu, Zeyuan Chen and colleagues at Salesforce AI Research introduce CLIFT, which trains a web agent to verify its own rollouts and reuses that verifier for test-time trajectory selection without an external judge.

49Agents
CUAWright: A Minimal Unified Interface for Digital Agents

CUAWright: A Minimal Unified Interface for Digital Agents

Yadong Lu and Ahmed Hassan Awadallah (Microsoft Research) with Yu Su, Huan Sun and colleagues at Ohio State present CUAWright, a harness in which computer-use agents act on web, desktop and CAD tasks only through terminal commands.

50Agents
ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning

ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning

Seil Kang, Hangoo Kang, Tarun Suresh and Azalia Mirhoseini (Stanford) with Yonsei, Korea University and Bespoke Labs present ThunderSyncRL, which streams gradient computation during agentic RL without introducing policy staleness.

51Agents
VERA: Scaling Verifiable Environments for Agentic co-Evolution

VERA: Scaling Verifiable Environments for Agentic co-Evolution

Junqi Liu, Yucheng Tang, Daguang Xu and colleagues at NVIDIA, with UC Santa Cruz, UIUC, NUS and Tsinghua, introduce VERA, which turns benchmark trajectories into 9,000+ verifiable sandboxes and uses them to co-evolve a model and its agent harness.

52Agents
What Does a Harness Buy? Tokens, Mostly

What Does a Harness Buy? Tokens, Mostly

Yangze Liu and Zhongyi Han (Shandong University) hold the model fixed and swap the harness across Claude Code, mini-SWE-agent and OpenCode on SWE-bench Verified, rerunning identical configurations to measure run-to-run noise.

53Agents
Correct Code, Broken Contributions? SWE-CC: Benchmarking Repository Policy Compliance for Coding Agents

Correct Code, Broken Contributions? SWE-CC: Benchmarking Repository Policy Compliance for Coding Agents

Hai Dang Truong, Rayner Goh and Yintong Huo (Singapore Management University) with Thanh Le-Cong (SUTD) introduce SWE-CC, a benchmark that checks whether coding agents follow a repository's contribution policies, not only whether their patches pass tests.

54Evaluation
TeleTune: Evolving Agent Skills From Offline Telemetry

TeleTune: Evolving Agent Skills From Offline Telemetry

Justin Chih-Yao Chen and Mohit Bansal (UNC Chapel Hill) with Elias Stengel-Eskin, Benjamin Van Durme, Gaurav Verma and colleagues at Microsoft introduce TeleTune, which learns a textual skill library for computer-use agents from offline user telemetry.

55Agents
Harness-Search: Guiding Long-Horizon Search through Multi-Agent Coordination

Harness-Search: Guiding Long-Horizon Search through Multi-Agent Coordination

Shanyong Wang, Yanyu Xu and colleagues at Xiaohongshu with Shandong University and UIUC introduce Harness-Search, which splits a long-horizon search agent into three permission-bounded roles that propose, commit and audit.

56Agents
Causal Improvement Graph for Agentic Harness Optimization

Causal Improvement Graph for Agentic Harness Optimization

Junjie Zhang, Shunyu Liu and Dacheng Tao (Nanyang Technological University) with Ting-En Lin and Yongbin Li (Tongyi Lab, Alibaba) introduce the Causal Improvement Graph (CIG), which stores the state of an automated harness-optimization search in a persistent graph instead of in the proposer's growing history.

57Agents
RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents

Mohamed Dhouib, Sonia Vanier and colleagues at École polytechnique (LIX), with Elie Bursztein of Google DeepMind, introduce RAISED, a self-distillation defense that trains a tool-using agent to behave on injected contexts exactly as it behaves on clean ones.

58Agents
Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild

Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild

Yifan Xiong and Yiling Lou (UIUC) with Jingyi Ge (UC Berkeley) and Zhenpeng Chen (Tsinghua) run the first empirical study of test adequacy in agent harnesses and introduce HarnessTester, a test generator built for harness code.

59Agents
D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?

D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?

Daifeng Li (HKUST and Alibaba) with Huiqiang Jiang, Dayiheng Liu and colleagues at Alibaba Group introduce D2K-Bench, which measures how well coding agents turn expert design guidance into efficient Triton GPU kernels.

60Agents
LEAP: Learning Efficient Action Proposals For LLM Agents

LEAP: Learning Efficient Action Proposals For LLM Agents

Zhen Xu, Qizheng Zhang, Gerry Wan, Shang Zhu and Ce Zhang (University of Chicago, Stanford, Together AI) present LEAP, which trains a 0.6B draft model to propose the next agent action so a larger target can verify it in parallel, cutting end-to-end agent latency.

61Agents
WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness

WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness

Yun-Yun Tsai (Columbia, during an internship at Meta) with Yuning Mao and colleagues at Meta Superintelligence Labs introduce WebUIProof, a benchmark that scores generated web interfaces by having a UI agent execute interaction tests in a headless browser.

62Evaluation
DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys and Peter W. J. Staar (ETH Zurich, IBM Research Zurich, Microsoft) build DeskForge, a controllable desktop environment that composes real applications and records dense annotations, and release the DeskForge-1M grounding corpus.

63Data
Coco: An Agentic Copilot for the Hardware--Software Co-Design Lifecycle

Coco: An Agentic Copilot for the Hardware--Software Co-Design Lifecycle

Samuel Kushnir, Amir Yazdanbakhsh, Parthasarathy Ranganathan, Suvinay Subramanian and colleagues at Google and Google DeepMind, with MIT, describe Coco, an agent platform deployed with TPU architects to set up hardware-software co-design experiments, run simulator sweeps and analyze the results.

64Agents
When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge

When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge

Xi Qin, Isabel Kurth and colleagues at SAP Lab build a meta-agent pipeline in which Claude Opus generates terminal tasks and verifiers for RL, and document why RL training of a 9B model still stalls.

65Agents
Trained Agentic Context Management

Trained Agentic Context Management

Bryce Sandlund, an independent researcher, fine-tunes Qwen3.6-35B-A3B to manage its own context through a two-tool harness (call itself with any prompt, read a token range of the input) and shows that an 8K-context model trained this way matches GPT-5.4 with a 1M context on long documents.

66Agents
Harness-Aware Distillation for Small Language Model Agents

Harness-Aware Distillation for Small Language Model Agents

Moonseok Choi, Taehong Moon, Giung Nam and Juho Lee (KAIST AI and an independent researcher) propose Harness-Aware Distillation (HAD), which distills a harness-equipped agent by teaching the student the decisions the teacher makes differently because of the harness.

67Training
When History Fails to Become Experience: Action Calibration in Language Agents

When History Fails to Become Experience: Action Calibration in Language Agents

Jingyu Liu and Yong Liu (Renmin University of China) with Zhiwen Wang, Yuxin Jing and Huanyu Zhou (ByteDance) study how language agents use their own interaction history and find that they rarely connect each action to its outcome.

68Agents
Continual Graph Memory for Mathematical Research Agents

Continual Graph Memory for Mathematical Research Agents

Junyi Zhang, Jinxi Yu, Eric Hanchen Jiang and colleagues at UCLA, with senior authors including Kai-Wei Chang, Raghu Meka, Nanyun Peng, Amit Sahai, Terence Tao and Wei Wang, present Ansatz, a mathematical research agent built around Continual Graph Memory, which stores proof progress as typed graphs instead of flat text.

69Memory
Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps

Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps

Jiaxin Zhang, Chien-Sheng Wu and colleagues at Salesforce AI Research propose Prospective Hindsight (PH), a reweighting rule that up-weights rollouts where the agent's own prediction of the outcome disagreed with the verifier's verdict; the paper is accepted at NeurIPS 2026.

70Reinforcement Learning
Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

Yu Li and colleagues at Southeast University introduce CITA, which trains a Comparative Inference Model to rank candidate next tool calls by their expected effect on final task success, and uses it to guide GRPO training and inference; the paper is a NeurIPS 2026 poster.

71Agents
Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents

Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents

Zhuowen Liu at the Japan Advanced Institute of Science and Technology re-evaluates fifteen prompt-injection detectors and two task-aware judges by replaying the ground-truth tool calls of AgentDojo and tau-bench, and finds that public benchmark scores do not predict detector behavior inside agents.

72Evaluation
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026