🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 6 – Sep 6, 2026
Efficiency · Evaluation

Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models

First page
Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models
The curator’s take

Ross Tieman and Evan Markou argue that semantic similarity is the wrong diversity measure for populations of language models, and use compression distance between raw outputs to recover the structure that predicts correlated failure.

Ask this paper

Key points
01

The distinction that matters: semantic similarity captures differences in the meaning of observed outputs, while generative-process diversity captures differences between the processes that could have produced them.

02

Normalised Compression Distance as the instrument, drawn from Algorithmic Information Theory and residualised against a permutation control so the measure is not just capturing output length or style.

03

Identifies population structure semantic similarity misses across 38 language models, which is the direct evidence that the two notions come apart.

04

Predicts correlated failure out of sample: cross-benchmark partial rank association of -0.216 with a 95% interval of [-0.309, -0.122], negative on all ten disjoint benchmark families, beyond semantic similarity and model-pair capability.

05

Why it matters: every ensemble, router, jury and multi-agent debate assumes its members fail independently. This gives a cheap measurable check on that assumption.

Abstract

Diversity is a widely observed factor in the resilient function of collective systems, yet the type of diversity that matters depends on the properties and failure modes of the system. This distinction is important for systems composed of multiple language models. Different models may be treated as independent components even when their behaviour and failures remain strongly correlated. Assessments of language-model populations using semantic similarity demonstrate limited semantic diversity, but this captures only differences in the meaning of observed outputs. We argue that a more fundamental notion of model diversity is generative-process diversity, the differences between processes capable of generating the observed outputs. Drawing from Algorithmic Information Theory, we use Normalised Compression Distance between raw model outputs, residualised against a permutation control, as a measure of inferred generative-process diversity. Across 38 language models, this measure identifies population structure missed by semantic similarity and predicts cross-task variation in chance-corrected correlated failure among model pairs across ten disjoint benchmark families, beyond semantic similarity and model-pair capability. The cross-benchmark partial rank association is $-0.216$ with a 95% interval of $[-0.309,-0.122]$, and the estimate is negative on all ten benchmarks. These results indicate that increased generative-process diversity is associated with reduced correlated failure in model pairs that is not attributable to semantic similarity or capability. Inferred generative-process diversity offers a novel and practical approach for investigating diversity of multi-model systems in safety-relevant contexts.

Every Monday
Get next week’s papers.
Subscribe on Substack