MagicVideo-V2

ByteDance's MagicVideo-V2 is an end-to-end text-to-video pipeline that stitches together four specialized modules into a high-fidelity generation system.
Ask this paper
Four-module pipeline: Combines a text-to-image model (T2I), a video motion generator, a reference-image embedding module, and a frame-interpolation module into a single flow.
High-resolution output: Produces high-resolution video with stronger motion fidelity and smoothness than prior T2V systems at comparable compute.
Reference conditioning: The reference-image embedding module lets the system align generated videos with stylistic or identity cues from a user-provided image.
User study wins: Reports preference wins over leading commercial T2V systems at the time on fidelity, motion quality, and prompt adherence in human evaluations.