🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 28, 2026
Agents · Evaluation

MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents

First page
MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents
The curator’s take

Chenxu Xiong, Mu Li, Alex Smola and colleagues at Boson AI introduce MSI-Bench, a benchmark for voice agents in conversations with several speakers, such as meetings and households.

Ask this paper

Key points
01

Test cases. 1,152 short multi-party audio scenes, split evenly between English and Mandarin, each with participant context, expected tool calls and atomic rubrics.

02

Three capability families. Multi-speaker memory, instruction following and reasoning. Averaged over 15 configurations, instruction following is easiest (48.9% and 46.7% all-rubrics pass) and memory is hardest (24.9% and 7.0%).

03

Top scores. The strongest configuration passes all rubrics on 66.8% of English cases (Gemini 3.1 Pro) and 54.5% of Mandarin cases. The best open-weight configuration reaches 34.0% and 19.3%.

04

Where open and frontier systems fail. Open-weight models are limited by the multi-speaker audio front end; Qwen2-Audio-7B returns valid output on every case and still passes 1.6%. Frontier systems fail at speaker-scoped decisions even on clean transcripts.

05

Unprompted responses. Models across the board often respond when nobody has addressed them.

Abstract

Voice provides a natural and immediate interface for AI agents. Many settings in which voice agents could be useful, including meetings, households, and collaborative work, are inherently multi-speaker. Supporting these settings introduces challenges that are largely absent from one-on-one interaction. We introduce the Multi-Speaker Interaction Benchmark (MSI-Bench) for evaluating multi-speaker voice interaction. Each test case is a short multi-party multi-turn audio scene with participant context, expected tool calls, and atomic rubrics. The benchmark targets three capability families: multi-speaker memory, multi-speaker instruction following, and multi-speaker reasoning. It comprises 1,152 test cases, evenly split between Mandarin Chinese and English (576 each). The strongest configuration on each split passes all rubrics on only 66.8% of English and 54.5% of Mandarin cases, and the strongest open-weight configuration on 34.0% and 19.3%. Failure analysis separates perception from reasoning: open-weight models are bottlenecked by the multi-speaker audio front-end, while frontier systems still fail speaker-scoped decision making on clean transcripts---and models across the board often respond when no one has addressed them. These results identify speaker-grounded perception, speaker-scoped decision making, and conversational restraint as concrete targets for future voice agents.

Every Monday
Get next week’s papers.
Subscribe on Substack