🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal · Robotics · Agents

Video Language Planning

First page
Video Language Planning
Paper summary

Enables synthesizing complex long-horizon video plans for robotics via tree search over vision-language and text-to-video models.

Ask this paper

Key points
01

Tree-search planner: Uses a tree-search procedure over a vision-language model serving as policy+value, with a text-to-video model acting as the dynamics model.

02

Long-horizon plans: Produces multi-step video plans for robotics tasks that would be infeasible with single-shot video generation.

03

Cross-domain generalization: Works across diverse robotics domains, showing the approach is not tied to a specific embodiment or task type.

04

Planning-via-generation: Demonstrates that generative video models can serve as world models for planning, a pattern that has gained traction through 2024.

Every Monday
Get next week’s papers.
Subscribe on Substack