🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

Sora

Paper preview
Sora
Paper summary

OpenAI unveils Sora, a text-to-video diffusion-transformer that generates coherent, minute-long 1080p videos from natural-language prompts.

Ask this paper

Key points
01

Spacetime patches: Videos are tokenized into spacetime patches and fed to a diffusion transformer, extending the scaling story of DiT-style image models to the video domain.

02

Minute-long generation: Produces videos up to 60 seconds with multiple characters, diverse motion types, and complex backgrounds while maintaining identity and scene consistency.

03

Multi-shot continuity: Can render multiple shots within a single video with persistence across characters and visual style, a capability prior text-to-video models struggled with.

04

World-simulator framing: OpenAI positions Sora as an early "world simulator" - still buggy on physics and object permanence, but a clear step toward generative models that encode intuitive physics.

Every Monday
Get next week’s papers.
Subscribe on Substack