Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving

Yi Xu, Chunqiang Tang and colleagues at Meta Platforms present Crossflow, a serving scheduler that lets decode nodes absorb prefill work when demand shifts, instead of relying on a fixed split between prefill and decode pools.
Ask this paper
Measured problem. In a large production LLM fleet the ratio of uncached input to output tokens has peak-to-mean ratios up to 4.7x at minute timescales. In a public agentic trace the hourly ratio varies by a median 24.5x within one day, while reassigning a replica takes tens of minutes.
Cost of static sizing. Provisioning each pool at its 95th percentile leaves up to 17% of cluster capacity idle; provisioning below it turns the imbalance into queueing.
Mechanism. Node roles never change. Each decode node publishes a short-lived, revocable lease that bounds local prefill compute, KV capacity, transfer work and projected output; the cluster scheduler reserves against it and a hard SLO violation sets the next lease to zero.
Results. Implemented in SGLang and run on NVIDIA GB300 servers with GPT-OSS-120B and GLM-5.2, Crossflow raises token throughput by 16.2% to 17.4% (geometric mean) over static P/D disaggregation, by up to 43.4% at high load, and lowers mean TTFT at every evaluated point.
Agentic relevance. Agent traffic, with its long cached prefixes and bursty tool-return prefills, is the case where the prefill/decode ratio swings most.
Abstract
As serving capacity demand surpasses that of training, serving efficiency becomes increasingly important. Prefill-decode (P/D) disaggregation improves serving efficiency through specialization and isolation of the two phases. These benefits rest on a static partitioning. Phase demand, however, is not static. We observe that in a large LLM fleet the ratio of uncached input to output tokens has peak-to-mean ratios up to 4.7x at minute timescales, and that in a public agentic trace the hourly ratio spans a median 24.5x within a single day, while reassigning a replica takes tens of minutes. Agentic traffic sharpens the mismatch. Sizing each pool at its ninety-fifth percentile leaves up to 17% of cluster capacity unused; sizing below it converts the same imbalance into queueing and unrealized throughput. We present Crossflow, which makes this boundary elastic without changing node roles. Each decode node publishes a short-lived, revocable lease that bounds local-prefill compute, KV capacity, transfer work, and projected output. Across public and internal traces, Crossflow improves token throughput by 16.2-17.4% on geometric mean over static P/D, and by up to 43.4% at high load, while reducing mean TTFT at every evaluated point.