Sora

OpenAI unveils Sora, a text-to-video diffusion-transformer that generates coherent, minute-long 1080p videos from natural-language prompts.
Ask this paper
Spacetime patches: Videos are tokenized into spacetime patches and fed to a diffusion transformer, extending the scaling story of DiT-style image models to the video domain.
Minute-long generation: Produces videos up to 60 seconds with multiple characters, diverse motion types, and complex backgrounds while maintaining identity and scene consistency.
Multi-shot continuity: Can render multiple shots within a single video with persistence across characters and visual style, a capability prior text-to-video models struggled with.
World-simulator framing: OpenAI positions Sora as an early "world simulator" - still buggy on physics and object permanence, but a clear step toward generative models that encode intuitive physics.