Skeleton-of-Thought (SoT)
Free while signed in. Answers cite the passages they came from.

Microsoft's Skeleton-of-Thought parallelizes LLM generation by first producing an answer skeleton then filling it in concurrently.
Two-stage generation: First generates an answer skeleton outlining the response structure, then fills in each skeleton point through parallel API calls.
2.39x speedup: Achieves up to 2.39x speedup over sequential decoding by exploiting the independence of skeleton points.
Quality improvements: Besides the speedup, reports quality improvements on some tasks - structure-first generation can produce more coherent long responses.
Applicability: Works best for list-style or outline-style responses where the skeleton decomposition is natural, less so for tightly coupled prose.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack