AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling
Yifan Zhang, Yutong Dai, Ran Xu, Zeyuan Chen and colleagues at Salesforce AI Research introduce CLIFT, which trains a web agent to verify its own rollouts and reuses that verifier for test-time trajectory selection without an external judge.

CUAWright: A Minimal Unified Interface for Digital Agents
Yadong Lu and Ahmed Hassan Awadallah (Microsoft Research) with Yu Su, Huan Sun and colleagues at Ohio State present CUAWright, a harness in which computer-use agents act on web, desktop and CAD tasks only through terminal commands.

ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning
Seil Kang, Hangoo Kang, Tarun Suresh and Azalia Mirhoseini (Stanford) with Yonsei, Korea University and Bespoke Labs present ThunderSyncRL, which streams gradient computation during agentic RL without introducing policy staleness.

VERA: Scaling Verifiable Environments for Agentic co-Evolution
Junqi Liu, Yucheng Tang, Daguang Xu and colleagues at NVIDIA, with UC Santa Cruz, UIUC, NUS and Tsinghua, introduce VERA, which turns benchmark trajectories into 9,000+ verifiable sandboxes and uses them to co-evolve a model and its agent harness.

What Does a Harness Buy? Tokens, Mostly
Yangze Liu and Zhongyi Han (Shandong University) hold the model fixed and swap the harness across Claude Code, mini-SWE-agent and OpenCode on SWE-bench Verified, rerunning identical configurations to measure run-to-run noise.

Correct Code, Broken Contributions? SWE-CC: Benchmarking Repository Policy Compliance for Coding Agents
Hai Dang Truong, Rayner Goh and Yintong Huo (Singapore Management University) with Thanh Le-Cong (SUTD) introduce SWE-CC, a benchmark that checks whether coding agents follow a repository's contribution policies, not only whether their patches pass tests.

TeleTune: Evolving Agent Skills From Offline Telemetry
Justin Chih-Yao Chen and Mohit Bansal (UNC Chapel Hill) with Elias Stengel-Eskin, Benjamin Van Durme, Gaurav Verma and colleagues at Microsoft introduce TeleTune, which learns a textual skill library for computer-use agents from offline user telemetry.

Harness-Search: Guiding Long-Horizon Search through Multi-Agent Coordination
Shanyong Wang, Yanyu Xu and colleagues at Xiaohongshu with Shandong University and UIUC introduce Harness-Search, which splits a long-horizon search agent into three permission-bounded roles that propose, commit and audit.

Causal Improvement Graph for Agentic Harness Optimization
Junjie Zhang, Shunyu Liu and Dacheng Tao (Nanyang Technological University) with Ting-En Lin and Yongbin Li (Tongyi Lab, Alibaba) introduce the Causal Improvement Graph (CIG), which stores the state of an automated harness-optimization search in a persistent graph instead of in the proposer's growing history.

RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
Mohamed Dhouib, Sonia Vanier and colleagues at École polytechnique (LIX), with Elie Bursztein of Google DeepMind, introduce RAISED, a self-distillation defense that trains a tool-using agent to behave on injected contexts exactly as it behaves on clean ones.

Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild
Yifan Xiong and Yiling Lou (UIUC) with Jingyi Ge (UC Berkeley) and Zhenpeng Chen (Tsinghua) run the first empirical study of test adequacy in agent harnesses and introduce HarnessTester, a test generator built for harness code.

D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?
Daifeng Li (HKUST and Alibaba) with Huiqiang Jiang, Dayiheng Liu and colleagues at Alibaba Group introduce D2K-Bench, which measures how well coding agents turn expert design guidance into efficient Triton GPU kernels.

LEAP: Learning Efficient Action Proposals For LLM Agents
Zhen Xu, Qizheng Zhang, Gerry Wan, Shang Zhu and Ce Zhang (University of Chicago, Stanford, Together AI) present LEAP, which trains a 0.6B draft model to propose the next agent action so a larger target can verify it in parallel, cutting end-to-end agent latency.

WebUIProof: Benchmarking WebUI Code Generators with UI-Agent Execution Harness
Yun-Yun Tsai (Columbia, during an internship at Meta) with Yuning Mao and colleagues at Meta Superintelligence Labs introduce WebUIProof, a benchmark that scores generated web interfaces by having a UI agent execute interaction tests in a headless browser.

DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents
A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys and Peter W. J. Staar (ETH Zurich, IBM Research Zurich, Microsoft) build DeskForge, a controllable desktop environment that composes real applications and records dense annotations, and release the DeskForge-1M grounding corpus.

Coco: An Agentic Copilot for the Hardware--Software Co-Design Lifecycle
Samuel Kushnir, Amir Yazdanbakhsh, Parthasarathy Ranganathan, Suvinay Subramanian and colleagues at Google and Google DeepMind, with MIT, describe Coco, an agent platform deployed with TPU architects to set up hardware-software co-design experiments, run simulator sweeps and analyze the results.

When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge
Xi Qin, Isabel Kurth and colleagues at SAP Lab build a meta-agent pipeline in which Claude Opus generates terminal tasks and verifiers for RL, and document why RL training of a 9B model still stalls.

Trained Agentic Context Management
Bryce Sandlund, an independent researcher, fine-tunes Qwen3.6-35B-A3B to manage its own context through a two-tool harness (call itself with any prompt, read a token range of the input) and shows that an 8K-context model trained this way matches GPT-5.4 with a 1M context on long documents.

Harness-Aware Distillation for Small Language Model Agents
Moonseok Choi, Taehong Moon, Giung Nam and Juho Lee (KAIST AI and an independent researcher) propose Harness-Aware Distillation (HAD), which distills a harness-equipped agent by teaching the student the decisions the teacher makes differently because of the harness.

When History Fails to Become Experience: Action Calibration in Language Agents
Jingyu Liu and Yong Liu (Renmin University of China) with Zhiwen Wang, Yuxin Jing and Huanyu Zhou (ByteDance) study how language agents use their own interaction history and find that they rarely connect each action to its outcome.

Continual Graph Memory for Mathematical Research Agents
Junyi Zhang, Jinxi Yu, Eric Hanchen Jiang and colleagues at UCLA, with senior authors including Kai-Wei Chang, Raghu Meka, Nanyun Peng, Amit Sahai, Terence Tao and Wei Wang, present Ansatz, a mathematical research agent built around Continual Graph Memory, which stores proof progress as typed graphs instead of flat text.

Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps
Jiaxin Zhang, Chien-Sheng Wu and colleagues at Salesforce AI Research propose Prospective Hindsight (PH), a reweighting rule that up-weights rollouts where the agent's own prediction of the outcome disagreed with the verifier's verdict; the paper is accepted at NeurIPS 2026.

Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents
Yu Li and colleagues at Southeast University introduce CITA, which trains a Comparative Inference Model to rank candidate next tool calls by their expected effect on final task success, and uses it to guide GRPO training and inference; the paper is a NeurIPS 2026 poster.

Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents
Zhuowen Liu at the Japan Advanced Institute of Science and Technology re-evaluates fifteen prompt-injection detectors and two task-aware judges by replaying the ground-truth tool calls of AgentDojo and tau-bench, and finds that public benchmark scores do not predict detector behavior inside agents.