Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit Signal

Yunxiang Mo, Donghao Zhao (HKUST) and Hejia Geng (University of Oxford) preregister a sweep of 3,520 self-consensus early-exit rules and find that none clears three acceptance gates, because agreement measures answer persistence rather than reasoning termination.
Ask this paper
The rule being tested: Repeatedly probe a partial reasoning trajectory for its current answer and stop once probes agree. The question is whether any such rule is both safe and token-saving, and whether one can be chosen once and reused.
The sweep and its result: 3,520 consensus rules replayed on frozen trajectories from two models and three benchmarks clear none of three acceptance gates fixed in advance. The frontier reproduces on a held-out split and on two unseen models, while a boundary-confidence control (DEER) swept through the same pipeline clears all three.
The consensus-termination gap: Agreement establishes that the current answer persists under a fixed probing procedure, not that reasoning has terminated. Stopping on it commits non-terminal answers.
The cost at a usable operating point: At a rule still saving 32% of tokens, one stop in nine fires on an answer the trajectory later abandons, and most of those cut off a correction the model would have made.
Widening the window does not fix it: The abandoned-answer share levels off near 7%, and by then the saving has fallen to 8%. Probe re-wording and a hand-labelled error taxonomy show the agreed answer is often a placeholder the model had not settled on.
Abstract
A natural way to cut reasoning-model inference cost is to repeatedly probe a single partial trajectory for its current answer and stop once probes agree -- self-consensus. We ask whether any such rule is both safe and token-saving, and whether one can be selected once and reused. A preregistered sweep of 3,520 consensus rules, replayed on frozen trajectories from two models and three benchmarks, clears none of three acceptance gates fixed in advance; the frontier reproduces on a held-out split and on two unseen models -- while a boundary-confidence control (DEER) swept through the same pipeline clears all three. The reason lies in the signal: agreement establishes that the current answer persists under a fixed probing procedure, not that the reasoning has terminated -- a consensus-termination gap. Stopping on it commits non-terminal answers. At a rule still saving 32% of the tokens, one stop in nine fires on an answer the trajectory itself later abandons, and most of those stops cut off a correction it would otherwise have made. Widening the agreement window does not remove them: the share levels off near 7%, and by then the saving has fallen to 8%. Probe re-wording and a hand-labelled error taxonomy show the agreed answer is often a placeholder the model had not settled on. Used on its own as the stop signal, agreement fails not because it is insufficiently strict, but because it repeatedly measures the wrong object.