🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

AudioGPT

First page
AudioGPT
Paper summary

Connects ChatGPT with audio foundational models for speech, music, sound, and talking head tasks.

Ask this paper

Key points
01

LLM as audio orchestrator: ChatGPT plans and dispatches audio tasks across specialist models (TTS, ASR, music generation, sound effects).

02

Modality transformation: Converts speech to text for ChatGPT processing, then generates speech from ChatGPT's text output.

03

Spoken dialogue: Enables end-to-end spoken dialogue where users talk to ChatGPT and it talks back.

04

Multi-modal agent pattern: An early example of the LLM-as-orchestrator pattern applied to audio, presaging 2024's fully multimodal voice agents.

Every Monday
Get next week’s papers.
Subscribe on Substack