TTSDS -- Text-to-Speech Distribution Score

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Minixhofer, Christoph, Klejch, Ondřej, Bell, Peter
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916502056206336
author Minixhofer, Christoph
Klejch, Ondřej
Bell, Peter
author_facet Minixhofer, Christoph
Klejch, Ondřej
Bell, Peter
contents Many recently published Text-to-Speech (TTS) systems produce audio close to real speech. However, TTS evaluation needs to be revisited to make sense of the results obtained with the new architectures, approaches and datasets. We propose evaluating the quality of synthetic speech as a combination of multiple factors such as prosody, speaker identity, and intelligibility. Our approach assesses how well synthetic speech mirrors real speech by obtaining correlates of each factor and measuring their distance from both real speech datasets and noise datasets. We benchmark 35 TTS systems developed between 2008 and 2024 and show that our score computed as an unweighted average of factors strongly correlates with the human evaluations from each time period.
format Preprint
id arxiv_https___arxiv_org_abs_2407_12707
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TTSDS -- Text-to-Speech Distribution Score
Minixhofer, Christoph
Klejch, Ondřej
Bell, Peter
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
Many recently published Text-to-Speech (TTS) systems produce audio close to real speech. However, TTS evaluation needs to be revisited to make sense of the results obtained with the new architectures, approaches and datasets. We propose evaluating the quality of synthetic speech as a combination of multiple factors such as prosody, speaker identity, and intelligibility. Our approach assesses how well synthetic speech mirrors real speech by obtaining correlates of each factor and measuring their distance from both real speech datasets and noise datasets. We benchmark 35 TTS systems developed between 2008 and 2024 and show that our score computed as an unweighted average of factors strongly correlates with the human evaluations from each time period.
title TTSDS -- Text-to-Speech Distribution Score
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2407.12707