🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 24, 2026
Agents · Data

When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain

First page
When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain
The curator’s take

Yunxiang Li, Xixin Wu and Helen Meng (CUHK) propose CIGAsk, an RL recipe that teaches a model both when to ask a clarifying question and how to phrase one that recovers the missing information (EMNLP 2026 Findings).

Ask this paper

Key points
01

Failure with prompting. Prompted models either ask for clarification on every query or ask vague questions that do not recover the missing detail.

02

How to ask. Counterfactual Information Gain compares the gold answer's log-likelihood under a frozen reference model with and without the user's reply, giving per-turn credit for informative questions.

03

When to ask. An Asymmetric Ambiguity Bonus gives a signed terminal reward based on whether the query was labeled ambiguous.

04

Training loop. Both rewards run inside multi-turn GRPO, with no separately trained critic.

05

Results. On three clarification benchmarks over table, passage and open-domain QA, CIGAsk-7B beats the strongest external baseline despite a smaller backbone, transfers across datasets without per-dataset tuning, and keeps single-turn QA performance.

Abstract

Instruction-tuned LLMs faced with underspecified queries often commit to a single interpretation rather than ask for clarification, producing confidently wrong answers. In our experiments, prompting alone is insufficient: models either ask for clarification on every query or ask vague questions that fail to recover the missing information. Addressing this failure requires learning two coupled skills: when to ask rather than answer and how to ask a question that recovers the disambiguating information. Existing recipes either address only one of these skills or require a separately trained critic. We propose CIGAsk, an RL recipe that teaches both skills through two complementary reward signals within a multi-turn GRPO loop. Counterfactual Information Gain (CIG) compares the gold-answer log-likelihood under a frozen reference model with and without the user response, providing per-turn credit that guides how to ask. The Asymmetric Ambiguity Bonus assigns a signed reward at the terminal token based on the gold ambiguity label, guiding when to ask. Across three clarification benchmarks spanning table, passage, and open-domain QA, CIGAsk-7B outperforms the strongest external baseline despite using a smaller backbone. It also transfers across datasets without per-dataset tuning while preserving single-turn QA performance on out-of-distribution benchmarks.

Every Monday
Get next week’s papers.
Subscribe on Substack