🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 3, 2026
Agents

Efficient Test-Time Adaptation through Human-AI Interaction

First page
Efficient Test-Time Adaptation through Human-AI Interaction
The curator’s take

Zora Zhiruo Wang, Apurva Gandhi, Rulin Shao and colleagues at Carnegie Mellon University, the University of Washington and Handshake propose TAHI, which turns the interaction history between one professional and their agent into both context and weight updates, plus an evolving per-user rubric that encodes the criteria the user never wrote down.

Ask this paper

Key points
01

The gap being closed is individual, not population-level: agents trained on population-scale data hit an average bar, while professional work requires the departure from average. TAHI treats cross-session interaction data as the supervision signal for that departure.

02

Two adaptation channels plus a rubric module: interaction signals go into agent context and into weights, and a rubric module accumulates each user's training and evaluation criteria as they surface during iterative editing.

03

Measured on 30 individuals, 600 tasks: across writing and visual creation, solo task success improves 4.5 to 20.9 percent within tens of tasks per user, so the adaptation budget is small.

04

The rubric doubles as an annotation tool: evolving rubrics catch 16.0 to 22.3 percent more failures than rubrics written by language models or by humans alone, which makes this an evaluation contribution as much as a personalization one.

05

Personalization transfers: agents adapted to one individual still improve success by up to 8.8 percent for other users, so the criteria learned per person are partly shared professional standards.

Abstract

AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet the artifacts they produce rarely meet the personal bar professionals need to stake their reputation on. On realistic, open-ended tasks where success criteria are heterogeneous and insufficiently documented, individual expertise lives precisely in the elevation and departure from the average. In practice, iterative human-agent interaction surfaces criteria that users cannot fully specify up front, yet apply repeatedly across tasks. We argue this cross-session interaction data is a rich, underused signal for closing the gap to individual expertise. In this work, we propose test-time adaptation through human-agent interaction (TAHI), which integrates these signals into agent context and weights, and crystallizes each user's training and evaluation criteria via an evolving rubric module. We adapt agents to 30 individuals in two high-utility domains, writing and visual creation, on a total of 600 tasks. Our agents improve solo task success by 4.5-20.9% within only tens of tasks. Meanwhile, our evolving rubric module serves as a scalable annotation tool, creating evaluation rubrics that catch 16.0-22.3% more failures than those from LMs or humans alone. While agents are adapted towards individuals, we show these personalized agents also produce improvements in success of up to 8.8% that generalize across users.

Every Monday
Get next week’s papers.
Subscribe on Substack