DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912650964762624 |
|---|---|
| author | Song, Yakun Zhuang, Xiaobin Chen, Jiawei Niu, Zhikang Yang, Guanrou Du, Chenpeng Jia, Dongya Chen, Zhuo Wang, Yuping Wang, Yuxuan Chen, Xie |
| author_facet | Song, Yakun Zhuang, Xiaobin Chen, Jiawei Niu, Zhikang Yang, Guanrou Du, Chenpeng Jia, Dongya Chen, Zhuo Wang, Yuping Wang, Yuxuan Chen, Xie |
| contents | Recent attempts to interleave autoregressive (AR) sketchers with diffusion-based refiners over continuous speech representations have shown promise, but they remain brittle under distribution shift and offer limited levers for controllability. We introduce DISTAR, a zero-shot text-to-speech framework that operates entirely in a discrete residual vector quantization (RVQ) code space and tightly couples an AR language model with a masked diffusion model, without forced alignment or a duration predictor. Concretely, DISTAR drafts block-level RVQ tokens with an AR language model and then performs parallel masked-diffusion infilling conditioned on the draft to complete the next block, yielding long-form synthesis with blockwise parallelism while mitigating classic AR exposure bias. The discrete code space affords explicit control at inference: DISTAR produces high-quality audio under both greedy and sample-based decoding using classifier-free guidance, supports trade-offs between robustness and diversity, and enables variable bit-rate and controllable computation via RVQ layer pruning at test time. Extensive experiments and ablations demonstrate that DISTAR surpasses state-of-the-art zero-shot TTS systems in robustness, naturalness, and speaker/style consistency, while maintaining rich output diversity. Audio samples are provided on https://anonymous.4open.science/w/DiSTAR_demo. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_12210 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation Song, Yakun Zhuang, Xiaobin Chen, Jiawei Niu, Zhikang Yang, Guanrou Du, Chenpeng Jia, Dongya Chen, Zhuo Wang, Yuping Wang, Yuxuan Chen, Xie Audio and Speech Processing Computation and Language Machine Learning Recent attempts to interleave autoregressive (AR) sketchers with diffusion-based refiners over continuous speech representations have shown promise, but they remain brittle under distribution shift and offer limited levers for controllability. We introduce DISTAR, a zero-shot text-to-speech framework that operates entirely in a discrete residual vector quantization (RVQ) code space and tightly couples an AR language model with a masked diffusion model, without forced alignment or a duration predictor. Concretely, DISTAR drafts block-level RVQ tokens with an AR language model and then performs parallel masked-diffusion infilling conditioned on the draft to complete the next block, yielding long-form synthesis with blockwise parallelism while mitigating classic AR exposure bias. The discrete code space affords explicit control at inference: DISTAR produces high-quality audio under both greedy and sample-based decoding using classifier-free guidance, supports trade-offs between robustness and diversity, and enables variable bit-rate and controllable computation via RVQ layer pruning at test time. Extensive experiments and ablations demonstrate that DISTAR surpasses state-of-the-art zero-shot TTS systems in robustness, naturalness, and speaker/style consistency, while maintaining rich output diversity. Audio samples are provided on https://anonymous.4open.science/w/DiSTAR_demo. |
| title | DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation |
| topic | Audio and Speech Processing Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2510.12210 |