🚀NEW LABGetting Started with Claude AgentsStart lab
Robotics

Genie

First page
Genie
Paper summary

DeepMind's Genie is an 11B-parameter foundation world model trained unsupervised on internet gameplay videos that generates action-controllable 2D worlds from a single image prompt.

Ask this paper

Key points
01

Three-component architecture: A spatiotemporal video tokenizer, an autoregressive dynamics model, and a scalable latent action model together learn to roll out playable worlds without any explicit action labels in training.

02

Prompt-conditioned worlds: Given a single image (photo, sketch, or video frame), Genie produces an interactive 2D environment the user can "play" frame-by-frame with inferred action controls.

03

Latent action space: The learned action embedding lets an agent imitate behaviors demonstrated in unseen videos, turning internet video into a generic source of demonstrations for training embodied policies.

04

Generalist-agent implication: Suggests a path toward training generalist agents entirely from observational video, without needing action-labeled datasets or curated simulators.

Every Monday
Get next week’s papers.
Subscribe on Substack