Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks

Reza Esfandiarpoor, Radek Osmulski, Even Oldridge and colleagues at NVIDIA (with the University of Edinburgh) measure what a ReAct retrieval agent adds over dense retrieval on complex search, and what it costs.
Ask this paper
Pipeline. An LLM runs a ReAct loop over a dense retriever with search, think and final-results tools. If the agent fails before answering, rankings from its earlier searches are merged with reciprocal rank fusion.
Accuracy. With the same embedding model, agentic retrieval improves nDCG@10 by 8.7 points over standard retrieval, and one pipeline is competitive on both the ViDoRe v3 and BRIGHT leaderboards, where specialized methods transfer poorly.
Cost. A query takes 107.4 seconds on average against 0.67 seconds for standard retrieval, and uses 764.1K input and 5.8K output tokens.
Infrastructure. Serving the retriever through MCP added a server per run and network latency per call; replacing it with an in-process, lock-protected singleton removed a class of deployment errors and raised GPU utilization. Code is in NVIDIA NeMo-Retriever.
Abstract
Modern information systems, including many agentic workflows, use dense retrieval to explore large amounts of unstructured data. However, dense retrieval relies on surface-level semantic similarity, which is insufficient for increasingly complex search applications. Here, we investigate agentic retrieval that combines the reasoning capabilities of Large Language Models (LLMs) with the efficient corpus exploration of retrievers in a ReAct agentic loop to solve complex retrieval tasks. In our experiments, we show that agentic retrieval is more effective than standard retrieval, improving nDCG@10 by 8.7 points using the same embedding model. Moreover, while specialized retrieval methods struggle on out-of-domain tasks, agentic retrieval is highly generalizable: the same pipeline achieves competitive results on both the ViDoRe v3 and BRIGHT leaderboards. However, this improvement comes at a cost. On average, agentic retrieval takes 107.4 seconds, compared to 0.67 seconds for standard retrieval, and consumes 764.1K input and 5.8K output tokens per query. In short, our study demonstrates the effectiveness of agentic retrieval in modern data systems and motivates future work on more cost-efficient retrieval agents for large-scale deployment.