BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kawamura, Masaya, Hasumi, Takuya, Shirahata, Yuma, Yamamoto, Ryuichi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910986230824960
author Kawamura, Masaya
Hasumi, Takuya
Shirahata, Yuma
Yamamoto, Ryuichi
author_facet Kawamura, Masaya
Hasumi, Takuya
Shirahata, Yuma
Yamamoto, Ryuichi
contents This paper proposes a highly compact, lightweight text-to-speech (TTS) model for on-device applications. To reduce the model size, the proposed model introduces two techniques. First, we introduce quantization-aware training (QAT), which quantizes model parameters during training to as low as 1.58-bit. In this case, most of 32-bit model parameters are quantized to ternary values {-1, 0, 1}. Second, we propose a method named weight indexing. In this method, we save a group of 1.58-bit weights as a single int8 index. This allows for efficient storage of model parameters, even on hardware that treats values in units of 8-bit. Experimental results demonstrate that the proposed method achieved 83 % reduction in model size, while outperforming the baseline of similar model size without quantization in synthesis quality.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03515
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
Kawamura, Masaya
Hasumi, Takuya
Shirahata, Yuma
Yamamoto, Ryuichi
Audio and Speech Processing
Machine Learning
Sound
Signal Processing
This paper proposes a highly compact, lightweight text-to-speech (TTS) model for on-device applications. To reduce the model size, the proposed model introduces two techniques. First, we introduce quantization-aware training (QAT), which quantizes model parameters during training to as low as 1.58-bit. In this case, most of 32-bit model parameters are quantized to ternary values {-1, 0, 1}. Second, we propose a method named weight indexing. In this method, we save a group of 1.58-bit weights as a single int8 index. This allows for efficient storage of model parameters, even on hardware that treats values in units of 8-bit. Experimental results demonstrate that the proposed method achieved 83 % reduction in model size, while outperforming the baseline of similar model size without quantization in synthesis quality.
title BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
topic Audio and Speech Processing
Machine Learning
Sound
Signal Processing
url https://arxiv.org/abs/2506.03515