🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Multimodal · Robotics · Agents

Video Language Planning

Free while signed in. Answers cite the passages they came from.

First page
Video Language Planning
The curator’s take

Enables synthesizing complex long-horizon video plans for robotics via tree search over vision-language and text-to-video models.

Key points
01

Tree-search planner: Uses a tree-search procedure over a vision-language model serving as policy+value, with a text-to-video model acting as the dynamics model.

02

Long-horizon plans: Produces multi-step video plans for robotics tasks that would be infeasible with single-shot video generation.

03

Cross-domain generalization: Works across diverse robotics domains, showing the approach is not tied to a specific embodiment or task type.

04

Planning-via-generation: Demonstrates that generative video models can serve as world models for planning, a pattern that has gained traction through 2024.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack