🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Reasoning · Safety

Metacognition in LLMs

First page
Metacognition in LLMs
Paper summary

Confidence calibration, self-verification, knowing when to stop, and knowing what you do not know have mostly been studied in isolation. This survey from Yale and UC Irvine argues they are facets of one capability, metacognition, and organizes the field around a monitor and control loop wrapped around the language model.

Ask this paper

Key points
01

Monitor and control framing: The model self-assesses before and after acting, then self-regulates by deciding whether to answer, retry, or defer, turning scattered behaviors into a single monitor-then-control cycle.

02

How it is measured: The survey catalogs psychology-based methods from signal detection theory, confidence-based metrics like calibration, AUROC, and ECE, activation-level neurofeedback, and interpretability probes such as concept injection.

03

How it is instilled and used: It reviews frameworks, architectures, prompting, and training that give LLMs, reasoning models, and agents metacognition, then shows gains in hallucination reduction, knowledge-boundary detection, and resistance to persuasion.

04

Why it matters: Metacognition underpins reliability, so a unified account of how to elicit, measure, and improve it gives builders a coherent target rather than a pile of one-off confidence tricks.

Every Monday
Get next week’s papers.
Subscribe on Substack