Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation

Yan Tang and colleagues formalize proactive service as a partially observable sequential decision process constrained by authorization and risk, where staying silent is a first-class action with option value.
Ask this paper
Four actions, not two: remain silent, ask, assist or act. Modeling silence and questions as decisions rather than absences is what lets timing be optimized at all.
Costs are named explicitly: interruption, misunderstanding, overreach and privacy, and most proactive-assistant work leaves these costs implicit.
One pipeline for the whole literature: state and need estimation, intervention gating, action construction and feedback adaptation, with prescribed, predictive, model-based and return-optimizing mechanisms as non-exclusive components.
Two claims worth arguing with: offline classification performance does not predict deployment benefit, and long-term memory is not a defining condition of proactivity, which contradicts a common product assumption.
What reliability actually requires: calibrated incremental intervention value, verifiable authorization, recoverable execution and counterfactual evidence.
Abstract
Large language model agents can plan, invoke tools, and modify external states, yet most systems still take an explicit user instruction as a fixed starting point. Proactive service moves the decision upstream: an agent must infer service opportunities from incomplete environmental and user signals, choose among remaining silent, asking, assisting, and acting, and account for interruption, misunderstanding, overreach, and privacy costs. This survey gives an operational definition centered on initiative and formulates the problem as a partially observable sequential decision process constrained by authorization and risk. The formulation represents timing, content, and delivery within one structured action, while making explicit the option value of waiting, the decision value of questions, and feedback-induced state changes. On this basis, we organize existing methods along one decision pipeline (state and need estimation, intervention gating, action construction, and feedback adaptation) and describe prescribed, predictive, model based, and return optimizing mechanisms as nonexclusive policy-construction components. We further normalize decision units and three-axis evidence descriptors across streaming dialogue, screen, video, software-engineering, and human-agent collaboration resources, and formalize metrics for triggering, timing, calibration, user burden, safety, and policy value. The synthesis shows why offline classification performance alone does not predict deployment benefit and why long-term memory is not a defining condition of proactivity. Reliable proactive service instead requires calibrated incremental intervention value, verifiable authorization, recoverable execution, and counterfactual evidence.