🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

Emu Video and Emu Edit

Paper preview
Emu Video and Emu Edit
Paper summary

Meta releases Emu Video and Emu Edit, a pair of diffusion models targeting controlled text-to-video generation and instruction-based image editing.

Ask this paper

Key points
01

Emu Video: Generates high-quality video from text-only, image-only, or combined text + image inputs using a factorized diffusion approach - text-to-image followed by image-conditioned video.

02

Emu Edit: Enables free-form image editing through text instructions, handling region, local, and global edits within one model.

03

Factorized video: The text-to-image then image-to-video split dramatically cuts training cost and improves controllability compared to end-to-end T2V models.

04

Unified research line: Both models extend Meta's Emu foundation family, pointing toward a unified multimodal generative stack shared across image, video, and edit tasks.

Every Monday
Get next week’s papers.
Subscribe on Substack