SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Hyeongju, Yang, Jinhyeok, Yu, Yechan, Ji, Seunghun, Morton, Jacob, Bous, Frederik, Byun, Joon, Lee, Juheon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918146078670848
author Kim, Hyeongju
Yang, Jinhyeok
Yu, Yechan
Ji, Seunghun
Morton, Jacob
Bous, Frederik
Byun, Joon
Lee, Juheon
author_facet Kim, Hyeongju
Yang, Jinhyeok
Yu, Yechan
Ji, Seunghun
Morton, Jacob
Bous, Frederik
Byun, Joon
Lee, Juheon
contents We introduce SupertonicTTS, a novel text-to-speech (TTS) system designed for efficient and streamlined speech synthesis. SupertonicTTS comprises three components: a speech autoencoder for continuous latent representation, a text-to-latent module leveraging flow-matching for text-to-latent mapping, and an utterance-level duration predictor. To enable a lightweight architecture, we employ a low-dimensional latent space, temporal compression of latents, and ConvNeXt blocks. The TTS pipeline is further simplified by operating directly on raw character-level text and employing cross-attention for text-speech alignment, thus eliminating the need for grapheme-to-phoneme (G2P) modules and external aligners. In addition, we propose context-sharing batch expansion that accelerates loss convergence and stabilizes text-speech alignment with minimal memory and I/O overhead. Experimental results demonstrate that SupertonicTTS delivers performance comparable to contemporary zero-shot TTS models with only 44M parameters, while significantly reducing architectural complexity and computational cost. Audio samples are available at: https://supertonictts.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23108
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
Kim, Hyeongju
Yang, Jinhyeok
Yu, Yechan
Ji, Seunghun
Morton, Jacob
Bous, Frederik
Byun, Joon
Lee, Juheon
Audio and Speech Processing
Machine Learning
Sound
We introduce SupertonicTTS, a novel text-to-speech (TTS) system designed for efficient and streamlined speech synthesis. SupertonicTTS comprises three components: a speech autoencoder for continuous latent representation, a text-to-latent module leveraging flow-matching for text-to-latent mapping, and an utterance-level duration predictor. To enable a lightweight architecture, we employ a low-dimensional latent space, temporal compression of latents, and ConvNeXt blocks. The TTS pipeline is further simplified by operating directly on raw character-level text and employing cross-attention for text-speech alignment, thus eliminating the need for grapheme-to-phoneme (G2P) modules and external aligners. In addition, we propose context-sharing batch expansion that accelerates loss convergence and stabilizes text-speech alignment with minimal memory and I/O overhead. Experimental results demonstrate that SupertonicTTS delivers performance comparable to contemporary zero-shot TTS models with only 44M parameters, while significantly reducing architectural complexity and computational cost. Audio samples are available at: https://supertonictts.github.io/.
title SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2503.23108