P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sach, Marvin, Fu, Yihui, Saijo, Kohei, Zhang, Wangyou, Cornell, Samuele, Scheibler, Robin, Li, Chenda, Kumar, Anurag, Wang, Wei, Qian, Yanmin, Watanabe, Shinji, Fingscheidt, Tim
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915408827645952
author Sach, Marvin
Fu, Yihui
Saijo, Kohei
Zhang, Wangyou
Cornell, Samuele
Scheibler, Robin
Li, Chenda
Kumar, Anurag
Wang, Wei
Qian, Yanmin
Watanabe, Shinji
Fingscheidt, Tim
author_facet Sach, Marvin
Fu, Yihui
Saijo, Kohei
Zhang, Wangyou
Cornell, Samuele
Scheibler, Robin
Li, Chenda
Kumar, Anurag
Wang, Wei
Qian, Yanmin
Watanabe, Shinji
Fingscheidt, Tim
contents In speech quality estimation for speech enhancement (SE) systems, subjective listening tests so far are considered as the gold standard. This should be even more true considering the large influx of new generative or hybrid methods into the field, revealing issues of some objective metrics. Efforts such as the Interspeech 2025 URGENT Speech Enhancement Challenge also involving non-English datasets add the aspect of multilinguality to the testing procedure. In this paper, we provide a brief recap of the ITU-T P.808 crowdsourced subjective listening test method. A first novel contribution is our proposed process of localizing both text and audio components of Naderi and Cutler's implementation of crowdsourced subjective absolute category rating (ACR) listening tests involving text-to-speech (TTS). Further, we provide surprising analyses of and insights into URGENT Challenge results, tackling the reliability of (P.808) ACR subjective testing as gold standard in the age of generative AI. Particularly, it seems that for generative SE methods, subjective (ACR MOS) and objective (DNSMOS, NISQA) reference-free metrics should be accompanied by objective phone fidelity metrics to reliably detect hallucinations. Finally, we will soon release our localization scripts and methods for easy deployment for new multilingual speech enhancement subjective evaluations according to ITU-T P.808.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11306
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge
Sach, Marvin
Fu, Yihui
Saijo, Kohei
Zhang, Wangyou
Cornell, Samuele
Scheibler, Robin
Li, Chenda
Kumar, Anurag
Wang, Wei
Qian, Yanmin
Watanabe, Shinji
Fingscheidt, Tim
Audio and Speech Processing
In speech quality estimation for speech enhancement (SE) systems, subjective listening tests so far are considered as the gold standard. This should be even more true considering the large influx of new generative or hybrid methods into the field, revealing issues of some objective metrics. Efforts such as the Interspeech 2025 URGENT Speech Enhancement Challenge also involving non-English datasets add the aspect of multilinguality to the testing procedure. In this paper, we provide a brief recap of the ITU-T P.808 crowdsourced subjective listening test method. A first novel contribution is our proposed process of localizing both text and audio components of Naderi and Cutler's implementation of crowdsourced subjective absolute category rating (ACR) listening tests involving text-to-speech (TTS). Further, we provide surprising analyses of and insights into URGENT Challenge results, tackling the reliability of (P.808) ACR subjective testing as gold standard in the age of generative AI. Particularly, it seems that for generative SE methods, subjective (ACR MOS) and objective (DNSMOS, NISQA) reference-free metrics should be accompanied by objective phone fidelity metrics to reliably detect hallucinations. Finally, we will soon release our localization scripts and methods for easy deployment for new multilingual speech enhancement subjective evaluations according to ITU-T P.808.
title P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge
topic Audio and Speech Processing
url https://arxiv.org/abs/2507.11306