DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving

Junyi Shen, Yao Lu and colleagues at the National University of Singapore introduce DynBranch, a serving layer that starts or reuses downstream agent work before a runtime branch decision has been made.
Ask this paper
Branch-resolution barrier. Agent workflows choose their execution path at runtime, so downstream work waits for the model or user to decide, even when that work is predictable or already computed.
Addressable branches. DynBranch gives an unresolved branch a stable coordinate, so candidate subgraphs can run speculatively and completed results can be cached and reused across requests.
Load-aware admission. A two-level controller only admits speculative work when its expected benefit exceeds the current load cost.
Results. Across four agent workloads with Qwen3-32B on 4x H200, mean latency falls by up to 32% against the strongest prior system per workload and 46-66% against no reuse, with unchanged workflow outputs.
Deployment. It runs at the model-API boundary with no changes to harnesses or inference engines, and the gains hold on Qwen3-8B on an RTX 4090.
Abstract
Agentic LLM workflows decide their execution paths at runtime. Downstream computation may be predictable, or may have run before, yet it cannot begin until the model or the user resolves the branch. We call this serialization the branch-resolution barrier. Caching alone does not hide it: the key that identifies a reusable result is not known until then. In this paper, we propose DynBranch, which makes an unresolved branch addressable before it resolves. Its stable coordinate lets candidate subgraphs run during resolution and completed subgraph results be reused across later requests. A two-level controller admits this work when its expected benefit exceeds the load price. DynBranch sits at the model-API boundary and requires no changes to agent harnesses or model execution engines. Across four agentic workloads with Qwen3-32B on 4x H200 GPUs, DynBranch reduces mean latency by up to 32% over each workload's strongest prior system and by 46-66% against a no-reuse floor, while preserving workflow results. The benefit persists across backbone families and on a commodity Qwen3-8B/RTX 4090 deployment.