🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

Translatotron 3

First page
Translatotron 3
Paper summary

Google's Translatotron 3 performs speech-to-speech translation using only monolingual data - no parallel corpora required.

Ask this paper

Key points
01

Fully unsupervised S2S: Learns direct speech-to-speech translation from monolingual data alone, a first for this task.

02

Three-component architecture: Combines a masked autoencoder for speech representation, unsupervised embedding mapping across languages, and back-translation for alignment.

03

Beats cascade baselines: Outperforms a comparable cascade of ASR + MT + TTS, a surprising result given cascade systems are typically the strong baseline.

04

Paralinguistic preservation: Preserves paralinguistic features - pauses, speaking rates, and speaker identity - that cascaded systems tend to wash out in translation.

Every Monday
Get next week’s papers.
Subscribe on Substack