ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

Zhenlong Dai, Jingyuan Chen and colleagues at Zhejiang University with Ant Group (NeurIPS 2026) introduce ToolSearcher, an RL framework for choosing tools from very large repositories through multi-turn search.
Ask this paper
The problem. Real tool repositories are too large to fit in context, and RL methods designed for knowledge QA do not handle distinguishing similar tools or composing compatible ones.
Three components. Category-constrained tool discrimination (trained first on subcategory-specific search) to separate functionally similar tools, event-level search modeling to reward discovering target tools during search, and trajectory-aligned credit allocation for the search and selection stages.
Results. On Qwen2.5-7B-Instruct, overall F1 rises from 9.8% to 51.3% and Match from 4.6% to 27.8% over the multi-turn baseline. On Qwen3-4B-Instruct, F1 rises from 40.8% to 53.1%.
Multi-tool scenarios. In the hardest intra-collection multi-tool setting (I3), ToolSearcher beats GDPO, MARAG-R1, GSPO and Search-R1 by 6.6 to 10.0 points.
Abstract
Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful tool use. Existing work often assumes a small or predefined set of tools, leaving large-scale tool selection underexplored. Real-world repositories contain a vast and diverse array of tools, making it difficult for LLMs to effectively search, distinguish, and compose tools under context-length constraints. We identify large-scale tool selection as a new challenge for agentic reinforcement learning, highlighting that existing RL methods for knowledge-based question answering are inadequate for selecting tools while considering compatibility. To address this challenge, we propose ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection. Specifically, we introduce category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools, event-level search modeling to explicitly optimize the discovery of target tools during multi-turn search, and trajectory-aligned credit allocation to provide fine-grained reward signals for different stages of the search-selection process. Extensive experiments on large-scale tool selection benchmarks demonstrate that ToolSearcher consistently outperforms a set of strong baselines in challenging settings involving iterative search and complex tool composition.