Follow the Entities: A Corpus Map for Agentic Search

Soyeong Jeong (KAIST, Microsoft intern), Sujay Kumar Jauhar, Sung Ju Hwang and Andrew Nam at Microsoft and KAIST build CorpusMap, an offline entity layer over a document collection that lets a search agent follow shared entities from one document to related ones instead of re-searching the flat corpus.
Ask this paper
Method. Mentions of the same entity are resolved across documents offline, and each entity gets an Entity Page that collects the facts stated about it with source attribution and links to every document that mentions it. The original documents stay available to the agent.
Results. Across 7 models and three benchmarks (EnterpriseRAG-Bench, WixQA, HERB), CorpusMap raises overall answer quality by 6.4 to 11.7 points over raw-corpus agentic search while using 34% to 57% fewer input tokens.
Against other navigation layers. It beats four alternatives, including an LLM Wiki in the style Karpathy described and Corpus2Skill.
Deployment. The map can be built without LLM calls and updated incrementally as documents are added, so the link structure is computed once and shared across queries.
Abstract
Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full corpus rather than reading only a fixed set of top-ranked documents. However, when the corpus is exposed only as a flat collection of files, a relevant document gives no indication of how it relates to others, so the agent must rediscover these relationships for every query, often missing complementary evidence while simultaneously consuming substantial additional tokens. To address this, we introduce CorpusMap, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources. Specifically, CorpusMap represents each recurring entity as an Entity Page that aggregates information about it and links to every document that refers to it, forming a graph between entities and documents that the agent can traverse to gather otherwise disconnected evidence. Moreover, since CorpusMap is constructed offline by resolving mentions of the same entity across documents, its links are shared across queries rather than rediscovered repeatedly at inference time. Using 7 different models with 3 benchmark datasets, we show that CorpusMap improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers, suggesting that entities serve as effective anchors for navigating large document collections.