VideoPoet
Free while signed in. Answers cite the passages they came from.

Google Research's VideoPoet is a large language model for zero-shot video generation that treats video as just another token stream.
Unified token stream: Uses multiple tokenizers to map video, image, audio, and text into a shared discrete token space for a single autoregressive model.
Zero-shot task variety: The same model handles image-to-video, video stylization, video-to-audio, and text-to-video without task-specific fine-tuning.
Language-model paradigm: Demonstrates that a plain autoregressive LM, given the right tokenizers, can handle video generation - challenging the diffusion-everywhere default for video.
Temporal consistency: Produces videos with reasonable motion coherence over short durations, a meaningful milestone for LM-based video generation.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack