🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 3, 2026
Evaluation

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

First page
Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
The curator’s take

Haoyuan Zhu at the University of Sheffield with Ranplan Wireless and Cambridge AI+ preregisters a reliability study of black-box LLM observers on shared serving endpoints and reports that the instrument itself is unstable enough to invalidate gates built on it.

Ask this paper

Key points
01

The core measurement: repeated calls to the same named model on a shared endpoint do not behave as a frozen instrument, so any threshold frozen on top of it inherits that drift.

02

Four attempted remedies all failed: waiting did not help across five further days (0.805 versus 0.800), switching providers did not help because four providers share the same floor with medians 0.74 to 0.88, provider metadata predicts none of it, and self-hosting on batch-invariant kernels helped only while the server was quiet.

03

Separation tracks error type, not error size: on constructed errors with known gaps the readout distinguishes kinds of error rather than magnitudes, which breaks the usual assumption behind scoring.

04

Deliverables for practitioners: a three-level snapshot-identity ladder, eight design rules and a reporting checklist.

05

A cheap prevention: a pilot at roughly 2 percent of the study's call volume would have exposed both unreachable gates in advance.

Abstract

Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99, each time with the execution record at ceiling. Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate gaps seven orders of magnitude below the instrument's own noise floor; and byte-identical inputs returning different rankings, a noise that exact-permutation readouts compound. Neither metric substitution nor sampling repaired it on the tested grid. Preregistered follow-ups bound the problem: waiting did not help on the days sampled (0.805 versus 0.800, replicated over five further days); switching providers did not help (four providers share the floor, medians 0.74 to 0.88, predicted by none of the metadata fields they expose); self-hosting on batch-invariant kernels helped only while the server was quiet; and on constructed errors with known gaps, the readout's separation tracks error type, not size. We distill the evidence into a three-level snapshot-identity ladder, eight design rules, and a reporting checklist; a pilot at roughly 2% of the study's call volume would have exposed both unreachable gates in advance. All results concern externally measured behaviour on shared serving infrastructure. On a shared endpoint, a model name is not a frozen instrument; a preregistered evaluation must measure its instrument before freezing any gate on it.

Every Monday
Get next week’s papers.
Subscribe on Substack