🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Evaluation · Agents · Reasoning

Can LLM Agents Infer World Models?

Free while signed in. Answers cite the passages they came from.

First page
Can LLM Agents Infer World Models?
The curator’s take

Can an LLM agent actually build a model of an environment it cannot see? This work makes that question gradeable through agentic automata learning. An agent has to uncover a hidden deterministic finite automaton by interacting with an oracle through two interfaces, membership queries that ask whether a string belongs to the target language, and equivalence queries that ask whether a proposed automaton is correct, which yields a clean, scalable testbed for interactive discovery.

Key points
01

A gradeable world-model test: Casting world-model inference as DFA learning gives objective success criteria and measurable interaction efficiency, with classic automata-learning algorithms as strong, well-understood baselines.

02

Controlled, scalable difficulty: The size of the hidden automaton acts as a difficulty knob, so the benchmark can scale task complexity smoothly rather than relying on a fixed set of puzzles.

03

Agents lag classic algorithms: Current agents can sometimes perform non-trivial interactive discovery, but performance drops sharply as DFA size grows, and trajectory analyses reveal recurring failures in query planning, evidence integration, and hypothesis construction.

04

Why it matters: Reasoning models clearly beat non-reasoning ones here, but the large gap to classic algorithms shows that systematic, interactive world-model building is still an unsolved capability rather than a byproduct of scale.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack