Translatotron 3
Free while signed in. Answers cite the passages they came from.

Google's Translatotron 3 performs speech-to-speech translation using only monolingual data - no parallel corpora required.
Fully unsupervised S2S: Learns direct speech-to-speech translation from monolingual data alone, a first for this task.
Three-component architecture: Combines a masked autoencoder for speech representation, unsupervised embedding mapping across languages, and back-translation for alignment.
Beats cascade baselines: Outperforms a comparable cascade of ASR + MT + TTS, a surprising result given cascade systems are typically the strong baseline.
Paralinguistic preservation: Preserves paralinguistic features - pauses, speaking rates, and speaker identity - that cascaded systems tend to wash out in translation.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack