🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

Moshi

First page
Moshi
Paper summary

introduces a speech-text foundation model and full-duplex spoken dialogue framework; they present several components of the systems; Helium is a 7B parameter text LLM; Mimi is a semantic-acoustic neural audio code with state-of-the-art performance on audio quality; a hierarchical multi-stream architecture that can generate arbitrary conversation in a speech-to-speech manner.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack