🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal · Data

LLaMa-Omni

First page
LLaMa-Omni
Paper summary

a model architecture for low-latency speech interaction with LLMs; it is based on Llama-3.1-8B-Instruct and can simultaneously generate both text and speech responses given speech instructions; responses can be generated with a response latency as low as 226ms; architecture-wise, it involves a speech encoder (Whispter-large-v3), a speech adaptor, an LLM, and a speech decoder; they also created a dataset of 200K speech interactions and responses.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack