🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Multimodal

Lumiere

Free while signed in. Answers cite the passages they came from.

First page
Lumiere
The curator’s take

Google's Lumiere is a space-time diffusion model for text-to-video that generates the entire video duration in a single forward pass rather than cascading short clips.

Key points
01

Space-Time U-Net (STUNet): A new architecture that jointly downsamples in both spatial and temporal dimensions, producing globally coherent motion instead of stitched-together keyframes.

02

Single-pass generation: Produces full-length videos in one pass, avoiding the motion discontinuities and temporal artifacts that plague cascaded models.

03

SoTA text-to-video: Achieves state-of-the-art results on standard text-to-video benchmarks in both quality and motion coherence.

04

Versatile video tasks: A single model supports image-to-video, stylized generation, cinemagraphs, and video inpainting - positioning Lumiere as a general-purpose video foundation model.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack