How Open is Open TTS? A Practical Evaluation of Open Source TTS Tools

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Răgman, Teodora, Stânea, Adrian Bogdan, Cucu, Horia, Stan, Adriana
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908930046689280
author Răgman, Teodora
Stânea, Adrian Bogdan
Cucu, Horia
Stan, Adriana
author_facet Răgman, Teodora
Stânea, Adrian Bogdan
Cucu, Horia
Stan, Adriana
contents Open-source text-to-speech (TTS) frameworks have emerged as highly adaptable platforms for developing speech synthesis systems across a wide range of languages. However, their applicability is not uniform -- particularly when the target language is under-resourced or when computational resources are constrained. In this study, we systematically assess the feasibility of building novel TTS models using four widely adopted open-source architectures: FastPitch, VITS, Grad-TTS, and Matcha-TTS. Our evaluation spans multiple dimensions, including qualitative aspects such as ease of installation, dataset preparation, and hardware requirements, as well as quantitative assessments of synthesis quality for Romanian. We employ both objective metrics and subjective listening tests to evaluate intelligibility, speaker similarity, and naturalness of the generated speech. The results reveal significant challenges in tool chain setup, data preprocessing, and computational efficiency, which can hinder adoption in low-resource contexts. By grounding the analysis in reproducible protocols and accessible evaluation criteria, this work aims to inform best practices and promote more inclusive, language-diverse TTS development. All information needed to reproduce this study (i.e. code and data) are available in our git repository: https://gitlab.com/opentts_ragman/OpenTTS
format Preprint
id arxiv_https___arxiv_org_abs_2603_24116
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Open is Open TTS? A Practical Evaluation of Open Source TTS Tools
Răgman, Teodora
Stânea, Adrian Bogdan
Cucu, Horia
Stan, Adriana
Audio and Speech Processing
Open-source text-to-speech (TTS) frameworks have emerged as highly adaptable platforms for developing speech synthesis systems across a wide range of languages. However, their applicability is not uniform -- particularly when the target language is under-resourced or when computational resources are constrained. In this study, we systematically assess the feasibility of building novel TTS models using four widely adopted open-source architectures: FastPitch, VITS, Grad-TTS, and Matcha-TTS. Our evaluation spans multiple dimensions, including qualitative aspects such as ease of installation, dataset preparation, and hardware requirements, as well as quantitative assessments of synthesis quality for Romanian. We employ both objective metrics and subjective listening tests to evaluate intelligibility, speaker similarity, and naturalness of the generated speech. The results reveal significant challenges in tool chain setup, data preprocessing, and computational efficiency, which can hinder adoption in low-resource contexts. By grounding the analysis in reproducible protocols and accessible evaluation criteria, this work aims to inform best practices and promote more inclusive, language-diverse TTS development. All information needed to reproduce this study (i.e. code and data) are available in our git repository: https://gitlab.com/opentts_ragman/OpenTTS
title How Open is Open TTS? A Practical Evaluation of Open Source TTS Tools
topic Audio and Speech Processing
url https://arxiv.org/abs/2603.24116