Qwen-Audio-Agent Technical Report

Alibaba Token Foundry presents Qwen-Audio-Agent, a harness that lets a full-duplex voice assistant keep talking while delegated tasks run in the background.
Ask this paper
Foreground and background. A Frontend Agent handles dialogue and chooses between calling a tool directly or delegating, and a Backend Agent carries out delegated work in its own context.
Orchestration runtime. The runtime tracks task state, coordinates requests for user input and authorization, and schedules when results return to the conversation.
Separated events. Speech interruption is kept separate from task cancellation, and task completion is kept separate from result delivery, so conversation continues during background work.
Results. On an in-house cockpit benchmark of 134 cases, mixed execution reaches 91.04% task success against 72.39% for direct-only and 80.60% for delegate-everything.
Latency. On matched successful turns, mixed execution lowers mean task latency by 26.73% and 30.91% relative to those two configurations.
Abstract
We present Qwen-Audio-Agent, a harness that combines full-duplex voice interaction with asynchronous task execution through a foreground-background architecture. A Frontend Agent manages dialogue and selects between direct tool use and delegation, while a Backend Agent carries out delegated tasks in a separate context. An Orchestration Runtime maintains task state, coordinates requests for user input and authorization, and schedules the return of results to the conversation. The runtime separates speech interruption from task cancellation and execution completion from result delivery, allowing conversation to continue while delegated work proceeds. Environmental events and persistent memory provide context within and across sessions. Independent adapters support integration with different frontend models, backend agents, and clients. We instantiate the architecture in desktop assistance, intelligent cockpits, and voice customer service. On an in-house cockpit benchmark of 134 cases, mixed execution achieves a task success rate of 91.04%, compared with 72.39% and 80.60% for the direct and all delegated configurations, respectively. In a separate latency evaluation on matched successful turns, mixed execution reduces mean task execution latency by 26.73% and 30.91% relative to these baselines, respectively. These results support the complementary use of direct tool calls for immediate operations and backend delegation for multi-step tasks.