VideoPoet

Google Research's VideoPoet is a large language model for zero-shot video generation that treats video as just another token stream.
Ask this paper
Unified token stream: Uses multiple tokenizers to map video, image, audio, and text into a shared discrete token space for a single autoregressive model.
Zero-shot task variety: The same model handles image-to-video, video stylization, video-to-audio, and text-to-video without task-specific fine-tuning.
Language-model paradigm: Demonstrates that a plain autoregressive LM, given the right tokenizers, can handle video generation - challenging the diffusion-everywhere default for video.
Temporal consistency: Produces videos with reasonable motion coherence over short durations, a meaningful milestone for LM-based video generation.