Instruction Duplication as an Inference-Time Control Primitive

Victor Lavrenko at PeaceTech VC measures instruction duplication, repeating only the procedural instruction, as a black-box inference-time control across seven instruction-tuned models and 16,800 scheduled generations, and reports gains on process compliance without any change in final-answer accuracy.
Ask this paper
Scope of the study: seven instruction-tuned models, 300 medical multiple-choice questions, eight placement conditions, 16,800 generations, with preregistered analysis.
The measured gain is procedural, not factual: the All-8 diagnostic, responses passing all eight observable tests, rises from 90.22 to 93.17 percent, removing 30.2 percent of the failures that survive one copy, while final-answer accuracy stays at exactly 60.21 percent.
A cost that has to be reported with it: premature commitment rises from 1.52 to 2.30 percent.
The author does not overclaim: a blinded audit gives 10 of 30 directional confirmations, 20 ties and no reversals, and its prespecified 28-of-30 criterion is not met.
Why the distinction matters operationally: when a downstream system repairs a generated trajectory, explicit trajectory state determines whether local repair is possible, which is where the procedural gain is spent.
Abstract
Procedural instruction following is a basic requirement for controllable language-model systems, especially when generated trajectories are inspected or repaired downstream. We introduce instruction duplication, a minimal black-box inference-time control that repeats only the procedural instruction, without retraining or decoding changes. Across seven instruction-tuned models, 300 medical multiple-choice questions, eight placement conditions, and 16,800 scheduled generations, moving from one to two copies raises the deterministic All-8 diagnostic--responses passing all eight observable tests--from 90.22% to 93.17% (+2.95 percentage points), eliminating 30.2% of the failures remaining after one copy. Pre-provisional TF-IDF recall rises from 73.44% to 74.81% (+1.38 points; Holm-adjusted p < .001), while final-answer accuracy remains exactly 60.21%. Premature commitment increases from 1.52% to 2.30% (p_Holm = .00536). A blinded challenge audit yields 10/30 directional confirmations, 20/30 perceptual ties, and no reversals; its prespecified 28/30 confirmation criterion is not met. Yet this distinction can matter operationally when a downstream system acts on the generated trajectory. In Answer Engineering (AE), where explicit trajectory state determines local repair, the published reason-first no-editing SSNHL endpoint was 25.1%; system-only AE was later reproduced at 84.2%, and the same trailing duplicate raised it to 97.1%. For conductive diagnostic branch preservation, the corresponding values are 58.9% published without editing, 78.6% with reproduced AE, and 73.8% with AE plus duplication--a within-AE decrease, but still 14.9 points above the no-editing baseline. Instruction duplication is therefore a low-complexity, placement-sensitive control whose practical value can emerge through the downstream system that consumes the exposed trajectory.