🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Multimodal

EMO: Emote Portrait Alive

Free while signed in. Answers cite the passages they came from.

First page
EMO: Emote Portrait Alive
The curator’s take

Alibaba's EMO synthesizes expressive talking-head videos directly from audio, bypassing the intermediate 3D models or facial landmarks used by prior approaches.

Key points
01

Audio2Video diffusion: A single diffusion model maps an audio waveform and a reference portrait directly to a video, letting subtle prosody drive facial motion without a handcrafted intermediate representation.

02

Speaking and singing: Produces convincing lip-sync and expressive facial motion for both speech and singing across varied styles, handling long-duration audio stably.

03

Identity preservation: Reports strong identity preservation and seamless frame transitions, outperforming prior methods on expressiveness and realism metrics.

04

Implications: Shows that eliminating the 3D/landmark intermediate stage - previously thought essential for controllable portraits - can actually improve fidelity and expressiveness for talking-head generation.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack