Voicebox
Paper preview

Paper summary
Meta's all-in-one generative speech model supporting 6 languages and many speech tasks in-context.
Ask this paper
01
Flow-matching training: Uses flow-matching with text-guided context to unify TTS, denoising, editing, and style transfer in one model.
02
20x faster: Outperforms specialized TTS systems while running 20x faster than prior state-of-the-art diffusion-based speech models.
03
Speech ICL: Supports in-context learning for speech - give it an audio prompt and it matches the speaker's voice, style, and prosody zero-shot.
04
Generalist speech: A major step toward generalist speech foundation models that would accelerate with 2024 systems like VoiceCraft and XTTS.