🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 16, 2026
Training

Large Language Models Develop Belief State Geometry In-Context

First page
Large Language Models Develop Belief State Geometry In-Context
The curator’s take

Daniel Balcells, Andrew Jun Lee, Chirag Rastogi, Adam Shai, Xavier Poncini and colleagues (Simplex, Astera Institute) test a computational-mechanics prediction: an LLM predicting hidden-Markov-model data in context should represent the belief state over hidden states.

Ask this paper

Key points
01

Setup: Six open LLMs are prompted with sequences from 40 HMMs chosen for non-trivial belief structure, and linear probes target the Bayesian posterior over hidden states.

02

Decodability: Belief states are linearly decodable from the residual stream with peak probe R^2 of 0.83 to 0.99, at layers ranging from early to late depending on the pair.

03

Causal role: Patching and steering the probe-identified subspace changes downstream predictions, so the representation is used and not only present.

04

Why it matters: The result gives a theory-derived, falsifiable target for what in-context learning computes at the level of latent state tracking.

Abstract

Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider such representations in a controlled setting: prompting LLMs with data emitted from hidden Markov models (HMMs) and probing for the corresponding belief state -- the posterior distribution over the HMM's hidden states given the observed token history. Across six open-source LLMs prompted with data from 40 HMMs selected for non-trivial belief structure, we find that belief states are linearly decodable from residual stream activations, with peak probe $R^2$-values from 0.83-0.99 across HMM and LLM combinations, ranging from early to late layers. To establish functional relevance, we intervene directly on the probe-identified subspace via patching and steering, resulting in downstream prediction quality on the order of the untampered model, while controls degrade performance substantially. Together, these results provide representation-level evidence that ICL in open-source LLMs approximates optimal Bayesian prediction over a context-inferred generative model. More broadly, our findings extend prior results linking input-distribution structure to activation geometry: from toy networks trained explicitly on HMM data to production-scale LLMs.

Every Monday
Get next week’s papers.
Subscribe on Substack