SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918146078670848 |
|---|---|
| author | Kim, Hyeongju Yang, Jinhyeok Yu, Yechan Ji, Seunghun Morton, Jacob Bous, Frederik Byun, Joon Lee, Juheon |
| author_facet | Kim, Hyeongju Yang, Jinhyeok Yu, Yechan Ji, Seunghun Morton, Jacob Bous, Frederik Byun, Joon Lee, Juheon |
| contents | We introduce SupertonicTTS, a novel text-to-speech (TTS) system designed for efficient and streamlined speech synthesis. SupertonicTTS comprises three components: a speech autoencoder for continuous latent representation, a text-to-latent module leveraging flow-matching for text-to-latent mapping, and an utterance-level duration predictor. To enable a lightweight architecture, we employ a low-dimensional latent space, temporal compression of latents, and ConvNeXt blocks. The TTS pipeline is further simplified by operating directly on raw character-level text and employing cross-attention for text-speech alignment, thus eliminating the need for grapheme-to-phoneme (G2P) modules and external aligners. In addition, we propose context-sharing batch expansion that accelerates loss convergence and stabilizes text-speech alignment with minimal memory and I/O overhead. Experimental results demonstrate that SupertonicTTS delivers performance comparable to contemporary zero-shot TTS models with only 44M parameters, while significantly reducing architectural complexity and computational cost. Audio samples are available at: https://supertonictts.github.io/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_23108 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System Kim, Hyeongju Yang, Jinhyeok Yu, Yechan Ji, Seunghun Morton, Jacob Bous, Frederik Byun, Joon Lee, Juheon Audio and Speech Processing Machine Learning Sound We introduce SupertonicTTS, a novel text-to-speech (TTS) system designed for efficient and streamlined speech synthesis. SupertonicTTS comprises three components: a speech autoencoder for continuous latent representation, a text-to-latent module leveraging flow-matching for text-to-latent mapping, and an utterance-level duration predictor. To enable a lightweight architecture, we employ a low-dimensional latent space, temporal compression of latents, and ConvNeXt blocks. The TTS pipeline is further simplified by operating directly on raw character-level text and employing cross-attention for text-speech alignment, thus eliminating the need for grapheme-to-phoneme (G2P) modules and external aligners. In addition, we propose context-sharing batch expansion that accelerates loss convergence and stabilizes text-speech alignment with minimal memory and I/O overhead. Experimental results demonstrate that SupertonicTTS delivers performance comparable to contemporary zero-shot TTS models with only 44M parameters, while significantly reducing architectural complexity and computational cost. Audio samples are available at: https://supertonictts.github.io/. |
| title | SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System |
| topic | Audio and Speech Processing Machine Learning Sound |
| url | https://arxiv.org/abs/2503.23108 |