Structured Output Collapses Diversity
Free while signed in. Answers cite the passages they came from.

Teams benchmark models in chat, then ship them behind JSON schemas for tools, extraction, and routing. This study of 44 language models shows that the structured surface you deploy is measurably more homogeneous than the chat surface you evaluated on.
JSON moves the defaults: Asking for JSON shifts 53% of a model's stable chat defaults, mostly back toward the crowd, and installs new defaults absent from chat, so the same model answers differently once wrapped in a schema.
Specific to trained formats: Diversity compression is significant for JSON and XML, absent for YAML and CSV, and reversed for an arbitrary bracket wrapper, which points to tool-use post-training rather than serialization itself as the cause.
Not the decoder: Enforcing the schema at the decoder compresses no further than simply requesting it, so the collapse lives in the model's response to the structured register, not in constrained decoding.
Why it matters: Diversity you measured in chat can vanish in production, quietly hurting sampling, synthetic data, and any workflow that depends on varied outputs, so structured surfaces deserve their own evaluation.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack