Video Language Planning
Free while signed in. Answers cite the passages they came from.

Enables synthesizing complex long-horizon video plans for robotics via tree search over vision-language and text-to-video models.
Tree-search planner: Uses a tree-search procedure over a vision-language model serving as policy+value, with a text-to-video model acting as the dynamics model.
Long-horizon plans: Produces multi-step video plans for robotics tasks that would be infeasible with single-shot video generation.
Cross-domain generalization: Works across diverse robotics domains, showing the approach is not tied to a specific embodiment or task type.
Planning-via-generation: Demonstrates that generative video models can serve as world models for planning, a pattern that has gained traction through 2024.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack