URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Ruiqi, Li, Xiquan, Chen, Wenxi, Niu, Zhikang, Yang, Chen, Ma, Ziyang, Yu, Kai, Chen, Xie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915435914461184
author Yan, Ruiqi
Li, Xiquan
Chen, Wenxi
Niu, Zhikang
Yang, Chen
Ma, Ziyang
Yu, Kai
Chen, Xie
author_facet Yan, Ruiqi
Li, Xiquan
Chen, Wenxi
Niu, Zhikang
Yang, Chen
Ma, Ziyang
Yu, Kai
Chen, Xie
contents Recent advances in large language models (LLMs) have driven significant progress in end-to-end spoken dialogue models (SDMs). In contrast to text-based LLMs, the evaluation framework for SDMs should encompass both cognitive dimensions (e.g., logical reasoning, knowledge) and speech-related aspects (e.g., paralinguistic cues, audio quality). However, there is still a lack of comprehensive evaluations for SDMs in speech-to-speech (S2S) scenarios. To address this gap, we propose URO-Bench, an extensive benchmark for SDMs. Notably, URO-Bench is the first S2S benchmark that covers evaluations about multilingualism, multi-round dialogues, and paralinguistics. Our benchmark is divided into two difficulty levels: basic track and pro track, each comprising 20 test sets, evaluating the spoken dialogue model's abilities in Understanding, Reasoning, and Oral conversation. Evaluations on our proposed benchmark reveal that current open-source SDMs perform rather well in daily QA tasks, but lag behind their backbone LLMs in terms of instruction-following ability and also suffer from catastrophic forgetting. Their performance in advanced evaluations of paralinguistic information and audio understanding remains subpar, highlighting the need for further research in this direction. We hope that URO-Bench can facilitate the development of spoken dialogue models by providing a multifaceted evaluation of existing models and helping to track progress in this area.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17810
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
Yan, Ruiqi
Li, Xiquan
Chen, Wenxi
Niu, Zhikang
Yang, Chen
Ma, Ziyang
Yu, Kai
Chen, Xie
Computation and Language
Audio and Speech Processing
Recent advances in large language models (LLMs) have driven significant progress in end-to-end spoken dialogue models (SDMs). In contrast to text-based LLMs, the evaluation framework for SDMs should encompass both cognitive dimensions (e.g., logical reasoning, knowledge) and speech-related aspects (e.g., paralinguistic cues, audio quality). However, there is still a lack of comprehensive evaluations for SDMs in speech-to-speech (S2S) scenarios. To address this gap, we propose URO-Bench, an extensive benchmark for SDMs. Notably, URO-Bench is the first S2S benchmark that covers evaluations about multilingualism, multi-round dialogues, and paralinguistics. Our benchmark is divided into two difficulty levels: basic track and pro track, each comprising 20 test sets, evaluating the spoken dialogue model's abilities in Understanding, Reasoning, and Oral conversation. Evaluations on our proposed benchmark reveal that current open-source SDMs perform rather well in daily QA tasks, but lag behind their backbone LLMs in terms of instruction-following ability and also suffer from catastrophic forgetting. Their performance in advanced evaluations of paralinguistic information and audio understanding remains subpar, highlighting the need for further research in this direction. We hope that URO-Bench can facilitate the development of spoken dialogue models by providing a multifaceted evaluation of existing models and helping to track progress in this area.
title URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2502.17810