Single response-rate metrics hide the real failures of multi-party voice assistants
on: Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue
Most spoken-dialogue benchmarks assume a single designated user and treat silence as a bug. Duplex-MPE treats silence as a correct answer and removes the designated user entirely. The benchmark places a voice assistant named Aria inside conversations among three or four humans, streams the full audio without transcripts or turn labels, and asks whether the assistant responds when addressed, stays quiet otherwise, and stops when a human resolves a pending request before Aria finishes answering.
The four scored capabilities are deliberately non-aggregable. Fresh-onset response rate counts only new decisions to speak after a request ends. Conditional answer accuracy is evaluated only when speech is present. Silence preservation scores every turn where the assistant should say nothing. Answering-window yield measures whether the assistant stops after a human resolves its pending request. Combining these into one number would hide the failure modes the benchmark is designed to expose.
The results show why that separation matters. Under explicit addressing, Freeze-Omni has the highest response presence yet the lowest fresh-onset response rate, because most of its coverage comes from speech already underway rather than new decisions to respond. FLM-Audio matches MiniCPM-o on response presence but its conditional answer accuracy is far lower—the source notes it produces correct answers on only a small share of response-present requests. Moshi has higher response presence than Voila yet much lower silence preservation. Each contrast would collapse under a single response-rate metric.
MiniCPM-o 4.5 leads on three of the four scored capabilities. The gap between its response presence and fresh-onset response rate is small, meaning its coverage reflects genuine new decisions. Its silence preservation and conditional answer accuracy both exceed the other four systems by statistically significant margins.
The paired explicit-versus-implicit design provides a useful control. A transcript-conditioned Gemini 3.1 Pro reference responds percentage points more often when the request names Aria than when the addressee must be inferred from context—a statistically significant paired difference. None of the five speech systems shows a significant paired difference in the same direction, which could mean they infer the implicit addressee successfully or that they are equally indiscriminate in both conditions. Response rates alone cannot distinguish these explanations.
The benchmark is fully synthetic: scripts generated by Claude Opus 5, speech synthesised by Qwen3-TTS, scoring by Claude Opus 5 and Qwen3-ASR-1.7B. Human reviewers validated sampled items and found high agreement with automated verdicts, but the scenarios do not capture real room acoustics, overlapping speech, or social dynamics. Results describe behaviour on these constructed scenes, not on deployed multi-party settings.
Four deliberately non-aggregable scores expose failure modes that a single response rate would bury in any multi-party spoken assistant evaluation.
Sources & links
Related on SkillFed
Benchmarking 34,198 real-world skills across three LLMs shows most of the benefit of agent skills disappears once agents must find their own -- pass rate lands within three points…
FORTIS benchmarks whether frontier LLM agents pick the minimally sufficient skill and stay inside its tool boundary. Across 2,143 queries, over-privilege is the default outcome…