AudioGPT
First page

Paper summary
Connects ChatGPT with audio foundational models for speech, music, sound, and talking head tasks.
Ask this paper
01
LLM as audio orchestrator: ChatGPT plans and dispatches audio tasks across specialist models (TTS, ASR, music generation, sound effects).
02
Modality transformation: Converts speech to text for ChatGPT processing, then generates speech from ChatGPT's text output.
03
Spoken dialogue: Enables end-to-end spoken dialogue where users talk to ChatGPT and it talks back.
04
Multi-modal agent pattern: An early example of the LLM-as-orchestrator pattern applied to audio, presaging 2024's fully multimodal voice agents.