Thinking Inertia: LLMs Keep Thinking When Told Not To

Dianqiao Lei (Tsinghua), Kevin Qinghong Lin and Philip Torr (Oxford), and Pan Lu and James Zou (Stanford) show that LLMs keep producing explicit reasoning when instructed not to, a behavior they call Thinking Inertia. Accepted at NeurIPS 2026.
Ask this paper
Measurement. Each response is split into a pre-answer trace and a final answer and scored at three levels: Empty-Thinking Rate for answer-only compliance, Question-Pre-answer Relevance, and an LLM-judge Explicit Inference Rate.
Proxies fail. Disabled thinking modes can still emit reasoning, and long traces can be filler, so mode settings and trace length are unreliable indicators of no-thinking.
Thinking Inertia. Across six prompting interventions on six LLMs, no-think controls do not reliably eliminate visible inference, and inference becomes more common as the answer space opens from Boolean to multiple-choice to open-ended.
Trade-off. Accuracy stays stable on Boolean and multiple-choice tasks, while open-ended tasks show a trade-off between answer-only compliance and accuracy. Human annotation confirms that even the strictest instruction leaves explicit inference on 57.1% of open-ended responses.
Answer space. Rewriting the same questions with candidate answers supplied makes answer-only responses easier to produce.
Abstract
Large Language Models (LLMs) increasingly ship with explicit "thinking modes", yet their counterpart, "no-thinking", has received far less attention. We study LLMs' no-thinking behavior along two axes. a. How to measure no-thinking? Prior work typically defines no-thinking through proxies such as a disabled thinking mode or the absence of long traces. These proxies are unreliable: disabled thinking modes may still emit reasoning, while long traces may contain filler rather than genuine inference. We instead normalize each response into a pre-answer trace and final answer, and evaluate it at three levels: (i) Empty-Thinking Rate for strict answer-only compliance; (ii) instruction-aware Question-Pre-answer Relevance for similarity between the question and pre-answer trace; and (iii) LLM-as-judge Explicit Inference Rate for visible explicit inference. Together, these metrics distinguish answer-only output, relevant but non-inferential text, and explicit inference. b. How does no-thinking vary across tasks and models? We evaluate six prompting interventions on six LLMs across Boolean, multiple-choice, and open-ended questions. We find that explicit no-think controls cannot reliably eliminate visible inference. Models instead exhibit "Thinking Inertia": explicit inference persists even under strict controls and becomes more prevalent as the answer space opens. Accuracy remains stable on Boolean and multiple-choice tasks, whereas open-ended tasks reveal a trade-off between answer-only compliance and task accuracy. Rewriting the same questions across answer spaces shows that supplying candidate answers makes answer-only responses easier to produce. These findings establish no-thinking as a non-trivial capability: stopping explicit reasoning cannot be assumed from model settings or instructions alone and deserves systematic evaluation alongside reasoning ability.