RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents

Mingxuan Zhang and colleagues present RAFT, which abstracts each closed support case into a directed chain of timeline entries and retrieves at the entry level rather than treating cases as static documents.
Ask this paper
Retrieval is anchored at a matched state. The system surfaces cases whose intermediate state matches the active case and returns the parent-case trajectory from that point, which is what makes the guidance actionable mid-case.
An optional case-level graph links cases. Through a configurable similarity representation, so related histories can be reached without re-embedding the whole corpus.
The retrieval layer is evaluated on its own. That requires no production deployment, which is the reason this result is reproducible where full agent evaluations are not.
Case Hit improves at every stage of case progress. Over vanilla RAG and GraphRAG, with statistically significant gains over the strongest baseline on a synthetic benchmark from Microsoft Learn Windows Server docs, and directional evidence from real Apache Jira duplicate labels.
Abstract
Effective troubleshooting agents in enterprise customer support depend on retrieving actionable guidance from similar historical cases, yet existing retrieval-augmented generation (RAG) systems treat support cases as static documents and overlook their multi-stage, stateful nature. We introduce RAFT (Retrieval-Augmented Framework for Troubleshooting Agents), a stateful RAG framework that abstracts each closed historical case into a directed chain of timeline entries and retrieves at the entry level, surfacing cases whose intermediate states match the active case and returning the parent-case trajectory anchored at the matched state; an optional case-level graph links cases through a configurable similarity representation. We evaluate this retrieval layer directly, which, unlike evaluating a full agent system, requires no production deployment. Because public multi-stage troubleshooting data is extremely rare, we pair a synthetic benchmark built from Microsoft Learn Windows Server documentation with real Apache Jira issues carrying human-created duplicate labels. RAFT improves Case Hit over vanilla RAG and GraphRAG baselines at every stage of case progress, with statistically significant gains over the strongest baseline; the Jira results provide directional evidence that the advantage transfers to real case histories. We release our benchmark, implementation, and the Apache Jira evaluation set.