🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

Audiobox

Paper preview
Audiobox
Paper summary

Meta's Audiobox is a unified flow-matching audio model that generates speech, sound effects, and music from natural-language and example prompts.

Ask this paper

Key points
01

Unified audio generation: Single model handles speech, sound, and music - ending the typical pattern of one model per audio modality.

02

Description + example prompting: Supports both natural-language descriptions and reference-audio examples for style control, letting users mix semantic and acoustic conditioning.

03

Self-supervised infilling: Adapts a self-supervised infilling objective to pretrain on large unlabeled audio, reducing dependence on scarce labeled speech/music datasets.

04

Novel voice/styles: Unlocks generation of novel vocal and acoustic styles by interpolating in the learned audio space, going beyond reproduction of training-set styles.

Every Monday
Get next week’s papers.
Subscribe on Substack