🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 10, 2026
Safety

When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination

First page
When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination
The curator’s take

Karan Parekh, Sanjana Pendyala Ravinder, Sana Mhapsekar and Medina Maloku (University of North Texas) plant 450 known contaminants across 150 academic papers and show that an LLM auditor's detection collapses as batch size grows, and that the failure mode at scale is fabrication rather than abstention.

Ask this paper

Key points
01

Detection falls off a cliff with batch size: Recovery of a 180-contaminant answer key across 60 documents is 50% on single documents, 60% on small batches, and 2.8% on large batches, using Google Gemini 3.0 Pro.

02

The model invents findings instead of reporting incomplete work: At large batch size the model produced confident findings including contaminants it invented, such as a telepathic squirrel and a quantum-powered toaster, which imitate the style of the planted material but appear in no document.

03

Plausible corruptions are the ones missed: Absurd insertions were recovered at 75% in completed evaluations, while semantic reversals and typographical corruptions were each recovered at only 50%. The corruptions most likely to occur in real documents are the ones detection misses.

04

Corpus: 150 academic papers spanning supply chain management and medical research, with 450 injected contaminants of three types, under three prompting regimes of increasing scale.

05

The harness the authors prescribe: Bounded batch sizes, direct content injection, and mechanical verification of every reported finding against source text.

Abstract

Large language models are increasingly proposed as automated auditors of document quality, yet their reliability as detectors of planted errors is poorly characterised. We construct a contaminated corpus of 150 academic papers spanning supply chain management and medical research, injecting 450 known contaminants of three types: typographical corruption, semantic reversal, and absurd out-of-context insertion. We then evaluate Google Gemini 3.0 Pro's ability to recover a 180-contaminant answer-key subset across 60 documents under three prompting regimes of increasing scale: single document, small batch, and large batch. Detection holds at small scale and then collapses: 50% recovery on single documents, 60% on small batches, and 2.8% on large batches. The failure mode at scale is not abstention but fabrication. Rather than reporting incomplete processing, the model produced confident findings including invented contaminants of its own, absurdities such as "telepathic squirrel" and "quantum-powered toaster" that mimic the style of the planted material but do not appear in any document. Detection also varies by contamination type: absurd insertions were recovered at 75% in completed evaluations, while semantic reversals and typographical corruptions were each recovered at only 50%. The corruptions most likely to occur in the wild, plausible ones, are the ones most often missed. We conclude that LLM document auditing degrades not gracefully but deceptively, and outline the harness such systems require: bounded batch sizes, direct content injection, and mechanical verification of every reported finding against source text.

Every Monday
Get next week’s papers.
Subscribe on Substack