Metacognition in LLMs
Free while signed in. Answers cite the passages they came from.

Confidence calibration, self-verification, knowing when to stop, and knowing what you do not know have mostly been studied in isolation. This survey from Yale and UC Irvine argues they are facets of one capability, metacognition, and organizes the field around a monitor and control loop wrapped around the language model.
Monitor and control framing: The model self-assesses before and after acting, then self-regulates by deciding whether to answer, retry, or defer, turning scattered behaviors into a single monitor-then-control cycle.
How it is measured: The survey catalogs psychology-based methods from signal detection theory, confidence-based metrics like calibration, AUROC, and ECE, activation-level neurofeedback, and interpretability probes such as concept injection.
How it is instilled and used: It reviews frameworks, architectures, prompting, and training that give LLMs, reasoning models, and agents metacognition, then shows gains in hallucination reduction, knowledge-boundary detection, and resistance to persuasion.
Why it matters: Metacognition underpins reliability, so a unified account of how to elicit, measure, and improve it gives builders a coherent target rather than a pile of one-off confidence tricks.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack