Compact Neural TTS Voices for Accessibility
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910803310936064 |
|---|---|
| author | Jain, Kunal Murphy, Eoin Gupta, Deepanshu Dyke, Jonathan Shah, Saumya Tsiaras, Vasilieios Petkov, Petko Conkie, Alistair |
| author_facet | Jain, Kunal Murphy, Eoin Gupta, Deepanshu Dyke, Jonathan Shah, Saumya Tsiaras, Vasilieios Petkov, Petko Conkie, Alistair |
| contents | Contemporary text-to-speech solutions for accessibility applications can typically be classified into two categories: (i) device-based statistical parametric speech synthesis (SPSS) or unit selection (USEL) and (ii) cloud-based neural TTS. SPSS and USEL offer low latency and low disk footprint at the expense of naturalness and audio quality. Cloud-based neural TTS systems provide significantly better audio quality and naturalness but regress in terms of latency and responsiveness, rendering these impractical for real-world applications. More recently, neural TTS models were made deployable to run on handheld devices. Nevertheless, latency remains higher than SPSS and USEL, while disk footprint prohibits pre-installation for multiple voices at once. In this work, we describe a high-quality compact neural TTS system achieving latency on the order of 15 ms with low disk footprint. The proposed solution is capable of running on low-power devices. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_17332 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Compact Neural TTS Voices for Accessibility Jain, Kunal Murphy, Eoin Gupta, Deepanshu Dyke, Jonathan Shah, Saumya Tsiaras, Vasilieios Petkov, Petko Conkie, Alistair Sound Machine Learning Audio and Speech Processing Contemporary text-to-speech solutions for accessibility applications can typically be classified into two categories: (i) device-based statistical parametric speech synthesis (SPSS) or unit selection (USEL) and (ii) cloud-based neural TTS. SPSS and USEL offer low latency and low disk footprint at the expense of naturalness and audio quality. Cloud-based neural TTS systems provide significantly better audio quality and naturalness but regress in terms of latency and responsiveness, rendering these impractical for real-world applications. More recently, neural TTS models were made deployable to run on handheld devices. Nevertheless, latency remains higher than SPSS and USEL, while disk footprint prohibits pre-installation for multiple voices at once. In this work, we describe a high-quality compact neural TTS system achieving latency on the order of 15 ms with low disk footprint. The proposed solution is capable of running on low-power devices. |
| title | Compact Neural TTS Voices for Accessibility |
| topic | Sound Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2501.17332 |