Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Varadhan, Praveen Srinivasa, Gulati, Amogh, Sankar, Ashwin, Anand, Srija, Gupta, Anirudh, Mukherjee, Anirudh, Marepally, Shiva Kumar, Bhatia, Ankur, Jaju, Saloni, Bhooshan, Suvrat, Khapra, Mitesh M.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912395740315648
author Varadhan, Praveen Srinivasa
Gulati, Amogh
Sankar, Ashwin
Anand, Srija
Gupta, Anirudh
Mukherjee, Anirudh
Marepally, Shiva Kumar
Bhatia, Ankur
Jaju, Saloni
Bhooshan, Suvrat
Khapra, Mitesh M.
author_facet Varadhan, Praveen Srinivasa
Gulati, Amogh
Sankar, Ashwin
Anand, Srija
Gupta, Anirudh
Mukherjee, Anirudh
Marepally, Shiva Kumar
Bhatia, Ankur
Jaju, Saloni
Bhooshan, Suvrat
Khapra, Mitesh M.
contents Despite rapid advancements in TTS models, a consistent and robust human evaluation framework is still lacking. For example, MOS tests fail to differentiate between similar models, and CMOS's pairwise comparisons are time-intensive. The MUSHRA test is a promising alternative for evaluating multiple TTS systems simultaneously, but in this work we show that its reliance on matching human reference speech unduly penalises the scores of modern TTS systems that can exceed human speech quality. More specifically, we conduct a comprehensive assessment of the MUSHRA test, focusing on its sensitivity to factors such as rater variability, listener fatigue, and reference bias. Based on our extensive evaluation involving 492 human listeners across Hindi and Tamil we identify two primary shortcomings: (i) reference-matching bias, where raters are unduly influenced by the human reference, and (ii) judgement ambiguity, arising from a lack of clear fine-grained guidelines. To address these issues, we propose two refined variants of the MUSHRA test. The first variant enables fairer ratings for synthesized samples that surpass human reference quality. The second variant reduces ambiguity, as indicated by the relatively lower variance across raters. By combining these approaches, we achieve both more reliable and more fine-grained assessments. We also release MANGO, a massive dataset of 246,000 human ratings, the first-of-its-kind collection for Indian languages, aiding in analyzing human preferences and developing automatic metrics for evaluating TTS systems.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12719
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation
Varadhan, Praveen Srinivasa
Gulati, Amogh
Sankar, Ashwin
Anand, Srija
Gupta, Anirudh
Mukherjee, Anirudh
Marepally, Shiva Kumar
Bhatia, Ankur
Jaju, Saloni
Bhooshan, Suvrat
Khapra, Mitesh M.
Computation and Language
Machine Learning
Sound
Audio and Speech Processing
Despite rapid advancements in TTS models, a consistent and robust human evaluation framework is still lacking. For example, MOS tests fail to differentiate between similar models, and CMOS's pairwise comparisons are time-intensive. The MUSHRA test is a promising alternative for evaluating multiple TTS systems simultaneously, but in this work we show that its reliance on matching human reference speech unduly penalises the scores of modern TTS systems that can exceed human speech quality. More specifically, we conduct a comprehensive assessment of the MUSHRA test, focusing on its sensitivity to factors such as rater variability, listener fatigue, and reference bias. Based on our extensive evaluation involving 492 human listeners across Hindi and Tamil we identify two primary shortcomings: (i) reference-matching bias, where raters are unduly influenced by the human reference, and (ii) judgement ambiguity, arising from a lack of clear fine-grained guidelines. To address these issues, we propose two refined variants of the MUSHRA test. The first variant enables fairer ratings for synthesized samples that surpass human reference quality. The second variant reduces ambiguity, as indicated by the relatively lower variance across raters. By combining these approaches, we achieve both more reliable and more fine-grained assessments. We also release MANGO, a massive dataset of 246,000 human ratings, the first-of-its-kind collection for Indian languages, aiding in analyzing human preferences and developing automatic metrics for evaluating TTS systems.
title Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation
topic Computation and Language
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2411.12719