Translatotron 3
First page

Paper summary
Google's Translatotron 3 performs speech-to-speech translation using only monolingual data - no parallel corpora required.
Ask this paper
01
Fully unsupervised S2S: Learns direct speech-to-speech translation from monolingual data alone, a first for this task.
02
Three-component architecture: Combines a masked autoencoder for speech representation, unsupervised embedding mapping across languages, and back-translation for alignment.
03
Beats cascade baselines: Outperforms a comparable cascade of ASR + MT + TTS, a surprising result given cascade systems are typically the strong baseline.
04
Paralinguistic preservation: Preserves paralinguistic features - pauses, speaking rates, and speaker identity - that cascaded systems tend to wash out in translation.