Spectron
First page

Paper summary
Google's Spectron is a spoken-language model trained end-to-end on raw spectrograms rather than text or discrete audio tokens.
Ask this paper
01
End-to-end spectrogram modeling: Processes spectrograms directly without an intermediate speech-recognition or tokenization step, preserving paralinguistic information.
02
High-quality spoken output: Fine-tuned to generate high-quality, accurate spoken language while preserving speaker and prosody characteristics.
03
Speaker preservation: Outperforms prior spoken-language models on speaker preservation - a known weakness of tokenizer-based approaches.
04
Semantic coherence: Also improves semantic coherence of generated speech, addressing the common drift problem in spectrogram-level generation.