Learn From Your Own Latents

LLMs learn by predicting tokens, while world models like JEPA and data2vec learn by predicting their own internal representations. This paper provides a sample-complexity theory for why the second approach can be dramatically more data-efficient, using a tractable probabilistic context-free grammar as the analytical setting where compositional structure can be measured exactly.
Ask this paper
Exponential gap in data efficiency: Predicting your own latents requires a number of samples that is constant in the tree depth L, whereas supervised and token-based self-supervised learning need samples that grow exponentially in L. The advantage is structural, not incidental.
Why latents win: Latent targets expose the compositional, hierarchical structure of the data directly, so the learner does not have to reconstruct it from surface tokens. That is the mechanism behind the data-efficiency gain.
Hierarchy may be implicit: The analysis suggests that explicit hierarchical stacking, as in H-JEPA, can be largely redundant, because methods like data2vec already learn hierarchical structure implicitly.
Why it matters: As token-prediction scaling laws press against data limits, this gives a principled argument for self-supervised objectives that predict abstractions instead of tokens. It is a theoretical foundation for why world-model-style training could beat brute-force next-token prediction on sample efficiency.