Toollery: Scaling LLM Agents to Thousands of Skills and Tools

Xiangxi Tian and Ran Guan (Huawei 2012 Laboratories) present Toollery, a training-free way to narrow libraries of thousands of skills and tools to a short candidate list before the LLM makes its final selection.
Ask this paper
Problem. Prompting with the full library grows token cost and latency with every added tool and adds distractors that lower selection accuracy.
Document-side query expansion. Toollery generates likely user-intent queries from each skill or tool specification and indexes them, so real requests match against intents rather than raw specs.
Skills and tools together. High-level skills and atomic tools are both treated as selectable capabilities, so one index serves skill libraries and tool registries.
Evaluation. Tested on the roughly 79K-capability SkillRouter benchmark, BFCL-V4 with over 440 tools, and 3,396 proprietary smart-cockpit requests over 220 tools.
Results. Recall beats plain specification retrieval, end-to-end selection improves on the cockpit data at top-10, and BFCL-V4 AST accuracy stays comparable; the authors note gains depend on workload coverage and provider caching.
Abstract
As LLM agents are exposed to hundreds to tens of thousands of skills, tools, and API functions, full-library prompting becomes costly, slow, and less reliable: each added candidate increases prompt tokens and latency, while longer candidate lists introduce more distractors for LLM selection. We present \textbf{Toollery}, a training-free candidate-compression framework for scalable LLM skill/tool selection. Following established document-side query expansion, Toollery generates user-intent queries from each skill/tool specification and builds a retrieval index that maps real user requests to compact candidate sets before final LLM decision-making. By treating high-level skills and atomic tools as selectable capabilities, Toollery can be applied to both skill libraries and tool registries. We evaluate Toollery on the roughly 79K-capability SkillRouter benchmark, BFCL-V4 with over 440 atomic tools, and 3,396 proprietary smart-cockpit requests over 220 tools. Across these settings, Toollery keeps online selection bounded to a compact top-$k$ candidate set and improves recall over ordinary specification retrieval. At a fixed top-10 budget, Toollery improves end-to-end selection on the cockpit dataset, and maintains comparable AST Accuracy on BFCL-V4. These results support Toollery as a practical candidate-compression framework for large and evolving agent capability libraries, while showing that quality and cost gains depend on workload coverage and provider caching.