🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Data · Training

Structured Output Collapses Diversity

First page
Structured Output Collapses Diversity
Paper summary

Teams benchmark models in chat, then ship them behind JSON schemas for tools, extraction, and routing. This study of 44 language models shows that the structured surface you deploy is measurably more homogeneous than the chat surface you evaluated on.

Ask this paper

Key points
01

JSON moves the defaults: Asking for JSON shifts 53% of a model's stable chat defaults, mostly back toward the crowd, and installs new defaults absent from chat, so the same model answers differently once wrapped in a schema.

02

Specific to trained formats: Diversity compression is significant for JSON and XML, absent for YAML and CSV, and reversed for an arbitrary bracket wrapper, which points to tool-use post-training rather than serialization itself as the cause.

03

Not the decoder: Enforcing the schema at the decoder compresses no further than simply requesting it, so the collapse lives in the model's response to the structured register, not in constrained decoding.

04

Why it matters: Diversity you measured in chat can vanish in production, quietly hurting sampling, synthetic data, and any workflow that depends on varied outputs, so structured surfaces deserve their own evaluation.

Every Monday
Get next week’s papers.
Subscribe on Substack