🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

Fugatto

First page
Fugatto
Paper summary

a new generative AI sound model (presented by NVIDIA) that can create and transform any combination of music, voices, and sounds using text and audio inputs, trained on 2.5B parameters and capable of novel audio generation like making trumpets bark or saxophones meow.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack