🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

AudioPaLM

First page
AudioPaLM
Paper summary

Fuses PaLM-2 and AudioLM into a multimodal architecture supporting speech understanding and generation.

Ask this paper

Key points
01

Unified speech-text: Represents both speech and text as tokens in a shared vocabulary, enabling any-to-any conversion between modalities.

02

Zero-shot translation: Performs zero-shot speech-to-text translation into languages never seen as translation targets during training.

03

Speech generation: Generates high-quality speech in the voice of the input speaker while preserving prosody.

04

Unified speech foundation: A precursor to 2024's fully multimodal systems like GPT-4o that natively process and generate speech.

Every Monday
Get next week’s papers.
Subscribe on Substack