🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 19, 2026
Agents

Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

First page
Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses
The curator’s take

Mahsa Amani and colleagues run the first end-to-end study of agentic Web search across ChatGPT, Claude, Grok and DeepSeek, combining real user interactions with controlled API experiments on the same models.

Ask this paper

Key points
01

Search invocation varies widely by platform. The decision of when to search at all differs substantially across platforms and models, and searching more often does not produce better responses.

02

Query formulation strategies differ. Agents use distinct complex querying strategies, so the retrieval behaviour is a property of the product, not only of the underlying model.

03

Platform search engines favour their own domains. The result sets returned show domain preferences tied to the platform rather than to the query.

04

Some claims cite nothing. Responses are largely grounded in retrieved results, but a portion of claims rest on uncited search results, which is a measurable attribution problem for anyone relying on these answers.

Abstract

Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSeek), combining real-world user interactions (invivo) with controlled experiments using the same platform's models by their APIs (invitro). We investigate the quality of agentic decisions to invoke Web search, their strategies to formulate queries, the potential domain preferences in the search results they receive, and the choices they make when transforming search results into grounded responses. We find that Web-search decisions vary substantially across platforms and models, while more frequent Web-search invocation does not necessarily yield better response quality. We further show that conversational agents employ different complex querying strategies and that platform specific search engines return search results from their preferred domains. Finally, although responses are largely grounded in search results, some claims rely on uncited search results, raising concerns about attribution and reliability. Our findings have important implications for the design of future AI agents and Web search tools optimized for conversational retrieval.

Every Monday
Get next week’s papers.
Subscribe on Substack