Evidence of Meaning in Language Models Trained on Programs
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsArgues LMs learn meaning despite only next-token prediction.
01
Programs as controlled input: Uses programs (which have well-defined semantics) to study whether LMs learn meaning versus surface patterns.
02
Intermediate-state prediction: Shows that LMs trained on programs learn to predict program state after each statement - evidence of semantic understanding.
03
Probe experiments: Careful probing experiments distinguish surface correlations from semantic representations.
04
Emergence argument: Adds empirical grounding to the "LLMs have world models" debate that dominated 2023's interpretability discussions.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack