Rational Clarification by Assistive Agents via Value-of-Information Reasoning

T. Duy Nguyen-Hien and Wee Sun Lee (NUS), Yee Whye Teh (Oxford) and Tan Zhi-Xuan introduce REVOIR, an inference-time method that decides whether an assistant should ask a clarifying question by estimating how much the answer would raise expected task reward, net of the cost of asking.
Ask this paper
Idea. Instead of asking until uncertainty about intent drops below a threshold, REVOIR computes the cost-adjusted value of information of each candidate question, building on assistance-game theory.
Results. On ambiguous QA (CondAmbigQA) and preference-aligned household planning (ADAPT), REVOIR gets higher success with fewer questions than prompting, chain-of-thought, fine-tuning or expected-information-gain baselines.
Against fine-tuning. On ADAPT it improves preference satisfaction by 13% to 15% over a fine-tuned clarification policy, with no training and five times fewer questions.
Adapts to cheap corrections. When the user can correct the assistant after it acts, REVOIR infers that asking is often not worth it.
Reasoning models ask less. Vanilla reasoning agents do not adapt their clarification behavior and ask fewer questions as reasoning effort increases.
Abstract
Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request --- risking misalignment with the user --- or ask a clarifying question. Which option is the most safe and helpful? A common approach is to ask questions that minimize uncertainty about the user's intent until a threshold is reached. However, this neglects the impact of uncertainty reduction on downstream performance, the costs of asking versus acting immediately, and the possibility that users may provide corrections without being asked. To navigate these trade-offs, we introduce Rational Enquiry via Value-of-Information Reasoning (REVOIR). REVOIR makes clarification decisions via inference-time reasoning about the value-of-information of a question, which captures the expected improvement in task reward due to the answer received. In two assistive tasks --- ambiguous question answering (CondAmbigQA) and preference-aligned household task planning (ADAPT) --- we show that REVOIR achieves greater success with fewer questions than approaches based on prompting, chain-of-thought, fine-tuning, or information gain, improving preference satisfaction on ADAPT by 13-15% over a fine-tuned clarification policy while requiring no training and asking five times fewer questions. Furthermore, when the assistant can receive cheap user corrections after acting, REVOIR naturally infers that asking questions is not always efficient, demonstrating the adaptivity of our approach. In contrast, we find that vanilla reasoning agents fail to adaptively clarify user requests, and request fewer clarifications as reasoning effort increases.