Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable

Pranav Aggarwal isolates a specific and trainable failure in LLM agents: the gate that decides whether to act at all collapses when evidence is packaged authoritatively, even when every number in the package is fabricated.
Ask this paper
The headline effect: Across 12 frontier models, commitment on a provably unpredictable question rises from 6.5% to 54.0% as evidence is escalated. A fully invented market panel lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data.
Three candidate explanations ruled out: Not incapacity: on matched answerable questions the same models answer almost always at near-perfect accuracy. Not belief: stated probabilities barely move across a gradient that swings action by 48 points. Not missing judgment: asked to classify knowability first, models call it irreducible 90% of the time and then commit on 0.4% of those.
The gate is separable and trainable: SFT of a 3B model on 540 synthetic cases, mostly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains.
And it is context-fragile: The gate holds only when the response format leaves room to reason. Rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly.
Why it matters: This is a clean dissociation between calibration and action. Agents that report good probabilities can still act on the unknowable, and the mechanism is the authority of packaging rather than information content.
Abstract
An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near-perfect accuracy. Nor is it belief - stated probabilities barely move across the gradient that swings action by 48 points, and score worse than a climatological baseline. Missing judgment isn't it either: asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails, and the effect is concentrated in a few models rather than universal. Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains. It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly. The gate is trainable and context-fragile, and deployment needs both halves of that sentence.