Evidence of Meaning in Language Models Trained on Programs
First page

Paper summary
Argues LMs learn meaning despite only next-token prediction.
Ask this paper
01
Programs as controlled input: Uses programs (which have well-defined semantics) to study whether LMs learn meaning versus surface patterns.
02
Intermediate-state prediction: Shows that LMs trained on programs learn to predict program state after each statement - evidence of semantic understanding.
03
Probe experiments: Careful probing experiments distinguish surface correlations from semantic representations.
04
Emergence argument: Adds empirical grounding to the "LLMs have world models" debate that dominated 2023's interpretability discussions.