🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

Voicebox

Paper preview
Voicebox
Paper summary

Meta's all-in-one generative speech model supporting 6 languages and many speech tasks in-context.

Ask this paper

Key points
01

Flow-matching training: Uses flow-matching with text-guided context to unify TTS, denoising, editing, and style transfer in one model.

02

20x faster: Outperforms specialized TTS systems while running 20x faster than prior state-of-the-art diffusion-based speech models.

03

Speech ICL: Supports in-context learning for speech - give it an audio prompt and it matches the speaker's voice, style, and prosody zero-shot.

04

Generalist speech: A major step toward generalist speech foundation models that would accelerate with 2024 systems like VoiceCraft and XTTS.

Every Monday
Get next week’s papers.
Subscribe on Substack