AudioPaLM
Free while signed in. Answers cite the passages they came from.

Fuses PaLM-2 and AudioLM into a multimodal architecture supporting speech understanding and generation.
Unified speech-text: Represents both speech and text as tokens in a shared vocabulary, enabling any-to-any conversion between modalities.
Zero-shot translation: Performs zero-shot speech-to-text translation into languages never seen as translation targets during training.
Speech generation: Generates high-quality speech in the voice of the input speaker while preserving prosody.
Unified speech foundation: A precursor to 2024's fully multimodal systems like GPT-4o that natively process and generate speech.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack