Skeleton-of-Thought (SoT)
First page

Paper summary
Microsoft's Skeleton-of-Thought parallelizes LLM generation by first producing an answer skeleton then filling it in concurrently.
Ask this paper
01
Two-stage generation: First generates an answer skeleton outlining the response structure, then fills in each skeleton point through parallel API calls.
02
2.39x speedup: Achieves up to 2.39x speedup over sequential decoding by exploiting the independence of skeleton points.
03
Quality improvements: Besides the speedup, reports quality improvements on some tasks - structure-first generation can produce more coherent long responses.
04
Applicability: Works best for list-style or outline-style responses where the skeleton decomposition is natural, less so for tightly coupled prose.