🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning · Efficiency

Skeleton-of-Thought (SoT)

First page
Skeleton-of-Thought (SoT)
Paper summary

Microsoft's Skeleton-of-Thought parallelizes LLM generation by first producing an answer skeleton then filling it in concurrently.

Ask this paper

Key points
01

Two-stage generation: First generates an answer skeleton outlining the response structure, then fills in each skeleton point through parallel API calls.

02

2.39x speedup: Achieves up to 2.39x speedup over sequential decoding by exploiting the independence of skeleton points.

03

Quality improvements: Besides the speedup, reports quality improvements on some tasks - structure-first generation can produce more coherent long responses.

04

Applicability: Works best for list-style or outline-style responses where the skeleton decomposition is natural, less so for tightly coupled prose.

Every Monday
Get next week’s papers.
Subscribe on Substack