MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Yuhao, Huang, Qianwei, Zhu, Guo, Dai, Zhanchen, Chen, Shunian, Zhu, Qiming, Pan, Le, Chen, Minghao, Zhang, Yuhao, Zhou, Li, Wang, Benyou, Li, Haizhou
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912587124310016
author Du, Yuhao
Huang, Qianwei
Zhu, Guo
Dai, Zhanchen
Chen, Shunian
Zhu, Qiming
Pan, Le
Chen, Minghao
Zhang, Yuhao
Zhou, Li
Wang, Benyou
Li, Haizhou
author_facet Du, Yuhao
Huang, Qianwei
Zhu, Guo
Dai, Zhanchen
Chen, Shunian
Zhu, Qiming
Pan, Le
Chen, Minghao
Zhang, Yuhao
Zhou, Li
Wang, Benyou
Li, Haizhou
contents The rapid advancement of speech-to-speech (S2S) large language models (LLMs) has significantly improved real-time spoken interaction. However, current evaluation frameworks remain inadequate for assessing performance in complex, multi-turn dialogues. To address this, we introduce MTalk-Bench, a multi-turn S2S benchmark covering three core dimensions: Semantic Information, Paralinguistic Information, and Ambient Sound. Each dimension includes nine realistic scenarios, along with targeted tasks to assess specific capabilities such as reasoning. Our dual-method evaluation framework combines Arena-style evaluation (pairwise comparison) and Rubrics-based evaluation (absolute scoring) for relative and absolute assessment. The benchmark includes both model and human outputs, evaluated by human evaluators and LLMs. Experimental results reveal two sets of findings. Overall performance of S2S LLMs: (1) models excel at semantic information processing yet underperform on paralinguistic information and ambient sounds perception; (2) models typically regain coherence by increasing response length, sacrificing efficiency in multi-turn dialogues; (3) modality-aware, task-specific designs outperform brute scaling. Evaluation framework and reliability: (1) Arena and Rubrics yield consistent, complementary rankings, but reliable distinctions emerge only when performance gaps are large; (2) LLM-as-a-judge aligns with humans when gaps are clear or criteria explicit, but exhibits position and length biases and is reliable on nonverbal evaluation only with text annotations. These results highlight current limitations in S2S evaluation and the need for more robust, speech-aware assessment frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18240
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
Du, Yuhao
Huang, Qianwei
Zhu, Guo
Dai, Zhanchen
Chen, Shunian
Zhu, Qiming
Pan, Le
Chen, Minghao
Zhang, Yuhao
Zhou, Li
Wang, Benyou
Li, Haizhou
Computation and Language
Artificial Intelligence
The rapid advancement of speech-to-speech (S2S) large language models (LLMs) has significantly improved real-time spoken interaction. However, current evaluation frameworks remain inadequate for assessing performance in complex, multi-turn dialogues. To address this, we introduce MTalk-Bench, a multi-turn S2S benchmark covering three core dimensions: Semantic Information, Paralinguistic Information, and Ambient Sound. Each dimension includes nine realistic scenarios, along with targeted tasks to assess specific capabilities such as reasoning. Our dual-method evaluation framework combines Arena-style evaluation (pairwise comparison) and Rubrics-based evaluation (absolute scoring) for relative and absolute assessment. The benchmark includes both model and human outputs, evaluated by human evaluators and LLMs. Experimental results reveal two sets of findings. Overall performance of S2S LLMs: (1) models excel at semantic information processing yet underperform on paralinguistic information and ambient sounds perception; (2) models typically regain coherence by increasing response length, sacrificing efficiency in multi-turn dialogues; (3) modality-aware, task-specific designs outperform brute scaling. Evaluation framework and reliability: (1) Arena and Rubrics yield consistent, complementary rankings, but reliable distinctions emerge only when performance gaps are large; (2) LLM-as-a-judge aligns with humans when gaps are clear or criteria explicit, but exhibits position and length biases and is reliable on nonverbal evaluation only with text annotations. These results highlight current limitations in S2S evaluation and the need for more robust, speech-aware assessment frameworks.
title MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.18240