Spectron
Free while signed in. Answers cite the passages they came from.

Google's Spectron is a spoken-language model trained end-to-end on raw spectrograms rather than text or discrete audio tokens.
End-to-end spectrogram modeling: Processes spectrograms directly without an intermediate speech-recognition or tokenization step, preserving paralinguistic information.
High-quality spoken output: Fine-tuned to generate high-quality, accurate spoken language while preserving speaker and prosody characteristics.
Speaker preservation: Outperforms prior spoken-language models on speaker preservation - a known weakness of tokenizer-based approaches.
Semantic coherence: Also improves semantic coherence of generated speech, addressing the common drift problem in spectrogram-level generation.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack