🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Robotics · Multimodal · Training

LingBot-World: Open-Source World Simulator

Free while signed in. Answers cite the passages they came from.

Figure 1
LingBot-World: Open-Source World Simulator
The curator’s take

LingBot-World is an open-source world simulator that evolves a video generation model into an interactive, real-time environment engine. Built on a 28B-parameter Mixture-of-Experts architecture, it achieves high-fidelity dynamics across diverse domains with sub-second latency at 16 fps, outperforming Genie 3 and Mirage 2 in dynamic degree while being fully open-source. - **Three-stage evolution pipeline:** A progressive training strategy transforms a pretrained video model into an interactive simulator: Stage I establishes a general video prior via the Wan2.2 14B model, Stage II injects world knowledge and action control through MoE middle-training on 60-second sequences, and Stage III adapts to causal attention with few-step distillation for real-time inference. - **Scalable data engine with hierarchical captioning:** A hybrid data engine ingests real-world footage, game engine recordings, and Unreal Engine synthetic data. A three-layer captioning strategy (narrative, scene-static, and dense temporal) disentangles motion control from scene generation, enabling precise action-contingent dynamics learning. - **Emergent spatial memory:** Without explicit 3D representations, the model maintains structural integrity of landmarks after 60 seconds out of view, reasons about unobserved state evolution (vehicles continuing trajectories off-screen), and supports coherent generation up to 10 minutes. VBench evaluation shows 0.8857 dynamic degree versus 0.76 for Yume-1.5 and 0.72 for HY-World 1.5. - **Versatile embodied AI applications:** Beyond visual synthesis, the framework supports promptable world events (global weather/style shifts and local object injection via text), an action agent trained on Qwen3-VL-2B for autonomous exploration, and 3D reconstruction from generated videos validating geometric consistency.

Key points
01

Three-stage evolution pipeline: A progressive training strategy transforms a pretrained video model into an interactive simulator: Stage I establishes a general video prior via the Wan2.2 14B model, Stage II injects world knowledge and action control through MoE middle-training on 60-second sequences, and Stage III adapts to causal attention with few-step distillation for real-time inference.

02

Scalable data engine with hierarchical captioning: A hybrid data engine ingests real-world footage, game engine recordings, and Unreal Engine synthetic data. A three-layer captioning strategy (narrative, scene-static, and dense temporal) disentangles motion control from scene generation, enabling precise action-contingent dynamics learning.

03

Emergent spatial memory: Without explicit 3D representations, the model maintains structural integrity of landmarks after 60 seconds out of view, reasons about unobserved state evolution (vehicles continuing trajectories off-screen), and supports coherent generation up to 10 minutes. VBench evaluation shows 0.8857 dynamic degree versus 0.76 for Yume-1.5 and 0.72 for HY-World 1.5.

04

Versatile embodied AI applications: Beyond visual synthesis, the framework supports promptable world events (global weather/style shifts and local object injection via text), an action agent trained on Qwen3-VL-2B for autonomous exploration, and 3D reconstruction from generated videos validating geometric consistency.

Abstract

We present LingBot-World, an open-sourced world simulator stemming from video generation. Positioned as a top-tier world model, LingBot-World offers the following features. (1) It maintains high fidelity and robust dynamics in a broad spectrum of environments, including realism, scientific contexts, cartoon styles, and beyond. (2) It enables a minute-level horizon while preserving contextual consistency over time, which is also known as "long-term memory". (3) It supports real-time interactivity, achieving a latency of under 1 second when producing 16 frames per second. We provide public access to the code and model in an effort to narrow the divide between open-source and closed-source technologies. We believe our release will empower the community with practical applications across areas like content creation, gaming, and robot learning.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack