🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal · Robotics · Efficiency

LingBot-World 2.0

First page
LingBot-World 2.0
Paper summary

Most world models fall apart after a few seconds, smearing textures and warping geometry as errors compound frame to frame. LingBot-World 2.0 from Robbyant holds 720p at 60 fps for a full hour of interaction and ships fully open.

Ask this paper

Key points
01

Causal backbone beats drift: A causal generation stack trained from the start to limit error accumulation replaces the usual bidirectional design, keeping scenes coherent well past the point where prior causal models collapse.

02

Durable teacher, real-time student: The high-capacity base model is distilled into a few-step student that renders in real time, so you get both long-horizon stability and responsive interaction from one system.

03

Act inside the world: Rather than only moving a camera, you can fight, draw a bow, cast spells, and type in events like weather changes, while an agentic harness of a scene-reading brain, a pilot, and a director keeps generating context-aware content.

04

Why it matters: Pairing hour-scale, real-time, high-fidelity generation with an open release, including a 14B model and a lighter single-GPU variant, gives researchers a serious interactive world model to build on rather than a closed demo.

Every Monday
Get next week’s papers.
Subscribe on Substack