BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910986230824960 |
|---|---|
| author | Kawamura, Masaya Hasumi, Takuya Shirahata, Yuma Yamamoto, Ryuichi |
| author_facet | Kawamura, Masaya Hasumi, Takuya Shirahata, Yuma Yamamoto, Ryuichi |
| contents | This paper proposes a highly compact, lightweight text-to-speech (TTS) model for on-device applications. To reduce the model size, the proposed model introduces two techniques. First, we introduce quantization-aware training (QAT), which quantizes model parameters during training to as low as 1.58-bit. In this case, most of 32-bit model parameters are quantized to ternary values {-1, 0, 1}. Second, we propose a method named weight indexing. In this method, we save a group of 1.58-bit weights as a single int8 index. This allows for efficient storage of model parameters, even on hardware that treats values in units of 8-bit. Experimental results demonstrate that the proposed method achieved 83 % reduction in model size, while outperforming the baseline of similar model size without quantization in synthesis quality. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_03515 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing Kawamura, Masaya Hasumi, Takuya Shirahata, Yuma Yamamoto, Ryuichi Audio and Speech Processing Machine Learning Sound Signal Processing This paper proposes a highly compact, lightweight text-to-speech (TTS) model for on-device applications. To reduce the model size, the proposed model introduces two techniques. First, we introduce quantization-aware training (QAT), which quantizes model parameters during training to as low as 1.58-bit. In this case, most of 32-bit model parameters are quantized to ternary values {-1, 0, 1}. Second, we propose a method named weight indexing. In this method, we save a group of 1.58-bit weights as a single int8 index. This allows for efficient storage of model parameters, even on hardware that treats values in units of 8-bit. Experimental results demonstrate that the proposed method achieved 83 % reduction in model size, while outperforming the baseline of similar model size without quantization in synthesis quality. |
| title | BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing |
| topic | Audio and Speech Processing Machine Learning Sound Signal Processing |
| url | https://arxiv.org/abs/2506.03515 |