🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 28, 2026
Agents

ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents

First page
ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents
The curator’s take

The ScholarSeed AI Team at Alibaba DAMO Academy introduces ScholarStack, which compiles a paper collection once into versioned, provenance-preserving research assets that scientific agents reuse across tasks.

Ask this paper

Key points
01

Three asset layers. Source-grounded paper-level statements, domain-level organization, and evidence-grounded cross-paper syntheses, all behind one access interface that returns the evidence granularity each task needs.

02

Evaluation. Four task families spanning ten settings, comparing agents with compiled assets against task-specific baselines on the same base models.

03

Gains on cross-paper tasks. Quality improves most where evidence spans papers, such as multi-paper QA and literature review generation (10.7% F1 improvement over matched raw-text access on one setting).

04

Token savings. On QASPER it reaches about 93% of full-text Answer-F1 with 34% of the input tokens, and on another setting about 93% of full-text Macro-F1 with about 11%. Answer generation uses 5.5k tokens per question versus 20.8k for full text.

05

Compile once. Assets are built once and reused, so query-time token cost falls on every task where it was measured.

Abstract

Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim assessment. Most existing systems, however, are organized around individual tasks: the same papers are repeatedly retrieved, segmented, and interpreted, and the understanding built in one task is difficult to reuse in the next. We present ScholarStack, a layered research asset framework that compiles a paper collection into reusable, versioned, and provenance-preserving assets at three complementary levels: source-grounded paper-level statements, domain-level organization, and evidence-grounded cross-paper syntheses. A common access interface returns task-specific views at the evidence granularity each task requires, preserving study conditions, source traceability, and verification status. We instantiate the framework on four task families spanning ten task settings, comparing agents that use the compiled assets with task-specific baselines under matched base models. Quality gains concentrate on tasks that require cross-paper evidence, such as multi-paper question answering and literature review generation, and query-time token cost falls on every task where it is measured, with assets compiled once and reused across tasks. These results suggest that layered research assets can serve as shared infrastructure for scientific agents, shifting literature-based assistance from isolated document processing toward cumulative, evidence-grounded workflows.

Every Monday
Get next week’s papers.
Subscribe on Substack