Compact Neural TTS Voices for Accessibility

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jain, Kunal, Murphy, Eoin, Gupta, Deepanshu, Dyke, Jonathan, Shah, Saumya, Tsiaras, Vasilieios, Petkov, Petko, Conkie, Alistair
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910803310936064
author Jain, Kunal
Murphy, Eoin
Gupta, Deepanshu
Dyke, Jonathan
Shah, Saumya
Tsiaras, Vasilieios
Petkov, Petko
Conkie, Alistair
author_facet Jain, Kunal
Murphy, Eoin
Gupta, Deepanshu
Dyke, Jonathan
Shah, Saumya
Tsiaras, Vasilieios
Petkov, Petko
Conkie, Alistair
contents Contemporary text-to-speech solutions for accessibility applications can typically be classified into two categories: (i) device-based statistical parametric speech synthesis (SPSS) or unit selection (USEL) and (ii) cloud-based neural TTS. SPSS and USEL offer low latency and low disk footprint at the expense of naturalness and audio quality. Cloud-based neural TTS systems provide significantly better audio quality and naturalness but regress in terms of latency and responsiveness, rendering these impractical for real-world applications. More recently, neural TTS models were made deployable to run on handheld devices. Nevertheless, latency remains higher than SPSS and USEL, while disk footprint prohibits pre-installation for multiple voices at once. In this work, we describe a high-quality compact neural TTS system achieving latency on the order of 15 ms with low disk footprint. The proposed solution is capable of running on low-power devices.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17332
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Compact Neural TTS Voices for Accessibility
Jain, Kunal
Murphy, Eoin
Gupta, Deepanshu
Dyke, Jonathan
Shah, Saumya
Tsiaras, Vasilieios
Petkov, Petko
Conkie, Alistair
Sound
Machine Learning
Audio and Speech Processing
Contemporary text-to-speech solutions for accessibility applications can typically be classified into two categories: (i) device-based statistical parametric speech synthesis (SPSS) or unit selection (USEL) and (ii) cloud-based neural TTS. SPSS and USEL offer low latency and low disk footprint at the expense of naturalness and audio quality. Cloud-based neural TTS systems provide significantly better audio quality and naturalness but regress in terms of latency and responsiveness, rendering these impractical for real-world applications. More recently, neural TTS models were made deployable to run on handheld devices. Nevertheless, latency remains higher than SPSS and USEL, while disk footprint prohibits pre-installation for multiple voices at once. In this work, we describe a high-quality compact neural TTS system achieving latency on the order of 15 ms with low disk footprint. The proposed solution is capable of running on low-power devices.
title Compact Neural TTS Voices for Accessibility
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2501.17332