The State Of TTS: A Case Study with Human Fooling Rates

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Varadhan, Praveen Srinivasa, Thomas, Sherry, S., Sai Teja M., Bhooshan, Suvrat, Khapra, Mitesh M.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913977247727616
author Varadhan, Praveen Srinivasa
Thomas, Sherry
S., Sai Teja M.
Bhooshan, Suvrat
Khapra, Mitesh M.
author_facet Varadhan, Praveen Srinivasa
Thomas, Sherry
S., Sai Teja M.
Bhooshan, Suvrat
Khapra, Mitesh M.
contents While subjective evaluations in recent years indicate rapid progress in TTS, can current TTS systems truly pass a human deception test in a Turing-like evaluation? We introduce Human Fooling Rate (HFR), a metric that directly measures how often machine-generated speech is mistaken for human. Our large-scale evaluation of open-source and commercial TTS models reveals critical insights: (i) CMOS-based claims of human parity often fail under deception testing, (ii) TTS progress should be benchmarked on datasets where human speech achieves high HFRs, as evaluating against monotonous or less expressive reference samples sets a low bar, (iii) Commercial models approach human deception in zero-shot settings, while open-source systems still struggle with natural conversational speech; (iv) Fine-tuning on high-quality data improves realism but does not fully bridge the gap. Our findings underscore the need for more realistic, human-centric evaluations alongside existing subjective tests.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04179
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The State Of TTS: A Case Study with Human Fooling Rates
Varadhan, Praveen Srinivasa
Thomas, Sherry
S., Sai Teja M.
Bhooshan, Suvrat
Khapra, Mitesh M.
Computation and Language
Machine Learning
Sound
Audio and Speech Processing
While subjective evaluations in recent years indicate rapid progress in TTS, can current TTS systems truly pass a human deception test in a Turing-like evaluation? We introduce Human Fooling Rate (HFR), a metric that directly measures how often machine-generated speech is mistaken for human. Our large-scale evaluation of open-source and commercial TTS models reveals critical insights: (i) CMOS-based claims of human parity often fail under deception testing, (ii) TTS progress should be benchmarked on datasets where human speech achieves high HFRs, as evaluating against monotonous or less expressive reference samples sets a low bar, (iii) Commercial models approach human deception in zero-shot settings, while open-source systems still struggle with natural conversational speech; (iv) Fine-tuning on high-quality data improves realism but does not fully bridge the gap. Our findings underscore the need for more realistic, human-centric evaluations alongside existing subjective tests.
title The State Of TTS: A Case Study with Human Fooling Rates
topic Computation and Language
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2508.04179