Position: Towards Responsible Evaluation for Text-to-Speech

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Yifan, Wang, Hui, Han, Bing, Liu, Shujie, Li, Jinyu, Qin, Yong, Chen, Xie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918521906135040
author Yang, Yifan
Wang, Hui
Han, Bing
Liu, Shujie
Li, Jinyu
Qin, Yong
Chen, Xie
author_facet Yang, Yifan
Wang, Hui
Han, Bing
Liu, Shujie
Li, Jinyu
Qin, Yong
Chen, Xie
contents Recent advances in text-to-speech (TTS) technology have enabled systems to generate speech that is often indistinguishable from human speech, bringing benefits to accessibility, content creation, and human-computer interaction. However, current evaluation practices are increasingly inadequate for capturing the full range of capabilities, limitations, and societal impacts of modern TTS systems. This position paper introduces the concept of Responsible Evaluation and argues that it is essential and urgent for the next phase of TTS development, structured through three progressive levels: (1) ensuring the faithful and accurate reflection of a model's true capabilities and limitations, with more robust, discriminative, and comprehensive objective and subjective scoring methodologies; (2) enabling comparability, standardization, and transferability through standardized benchmarks, transparent reporting, and transferable evaluation metrics; and (3) assessing governance, fairness, and security concerns around data provenance, disparities, misuse, spoofing, and traceability. Through this concept, we critically examine current evaluation practices, identify systemic shortcomings, and propose actionable recommendations. We hope this concept will not only foster more reliable TTS technology but also guide its development toward ethically sound and societally beneficial applications.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06927
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Position: Towards Responsible Evaluation for Text-to-Speech
Yang, Yifan
Wang, Hui
Han, Bing
Liu, Shujie
Li, Jinyu
Qin, Yong
Chen, Xie
Audio and Speech Processing
Recent advances in text-to-speech (TTS) technology have enabled systems to generate speech that is often indistinguishable from human speech, bringing benefits to accessibility, content creation, and human-computer interaction. However, current evaluation practices are increasingly inadequate for capturing the full range of capabilities, limitations, and societal impacts of modern TTS systems. This position paper introduces the concept of Responsible Evaluation and argues that it is essential and urgent for the next phase of TTS development, structured through three progressive levels: (1) ensuring the faithful and accurate reflection of a model's true capabilities and limitations, with more robust, discriminative, and comprehensive objective and subjective scoring methodologies; (2) enabling comparability, standardization, and transferability through standardized benchmarks, transparent reporting, and transferable evaluation metrics; and (3) assessing governance, fairness, and security concerns around data provenance, disparities, misuse, spoofing, and traceability. Through this concept, we critically examine current evaluation practices, identify systemic shortcomings, and propose actionable recommendations. We hope this concept will not only foster more reliable TTS technology but also guide its development toward ethically sound and societally beneficial applications.
title Position: Towards Responsible Evaluation for Text-to-Speech
topic Audio and Speech Processing
url https://arxiv.org/abs/2510.06927