Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

Ido Levy, Asaf Yehudai, Segev Shlomov, Asaf Adi and Leshem Choshen at IBM Research and the Weizmann Institute study which unrequested information an agent should go after, and train an 8B questioner to choose the questions that retrieve the most required evidence.
Ask this paper
Two kinds of proactivity. Horizontal proactivity pursues unstated information the current context already points to; vertical proactivity pursues needs that only earlier evidence reveals. A need graph recovered from each benchmark's own decomposition scores both, plus whether the agent stops at the right time, without a model judge.
Q&D method. A questioner is trained to prefer the question whose continuation retrieves more of the required evidence. There is no reward model and no judge.
Results. On held-out splits of three multi-hop QA benchmarks at equal retrieval spend, the trained Qwen3-8B questioner improves both forms of proactivity and beats a prompted model 15 times larger on two of the three, with the gain holding after controlling for question count and length.
Transfer to customer service. Placed without further training into an interactive agent with a simulated customer, it completes more tasks while asking fewer questions, and in retail it beats the 15x larger model with fewer customer follow-up turns.
Caveats the authors state. One round of off-policy training, small retrieval pools, and every user-facing number comes from a simulated customer.
Abstract
An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. A need graph, recovered from a benchmark's own decomposition, records which needs depend on which, so both forms, and whether the agent stops at the right time, can be scored from a transcript without a model judge. To learn this behavior, we propose Q&D (questioner and drafter), which trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge. On held-out splits of three multi-hop question-answering benchmarks, at equal retrieval spend, the trained questioner improves both forms of proactivity over the same model, prompted, and outperforms a prompted model $15\times$ larger in the same role on two of the three, and the gain persists after controlling for question volume and length. Without further training, we place the questioner in an interactive customer-service agent with a simulated customer, where it completes more tasks while asking fewer questions, and in retail it outperforms the $15\times$ larger model with fewer follow-up turns from the customer. These results show that proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.