DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Yakun, Zhuang, Xiaobin, Chen, Jiawei, Niu, Zhikang, Yang, Guanrou, Du, Chenpeng, Jia, Dongya, Chen, Zhuo, Wang, Yuping, Wang, Yuxuan, Chen, Xie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912650964762624
author Song, Yakun
Zhuang, Xiaobin
Chen, Jiawei
Niu, Zhikang
Yang, Guanrou
Du, Chenpeng
Jia, Dongya
Chen, Zhuo
Wang, Yuping
Wang, Yuxuan
Chen, Xie
author_facet Song, Yakun
Zhuang, Xiaobin
Chen, Jiawei
Niu, Zhikang
Yang, Guanrou
Du, Chenpeng
Jia, Dongya
Chen, Zhuo
Wang, Yuping
Wang, Yuxuan
Chen, Xie
contents Recent attempts to interleave autoregressive (AR) sketchers with diffusion-based refiners over continuous speech representations have shown promise, but they remain brittle under distribution shift and offer limited levers for controllability. We introduce DISTAR, a zero-shot text-to-speech framework that operates entirely in a discrete residual vector quantization (RVQ) code space and tightly couples an AR language model with a masked diffusion model, without forced alignment or a duration predictor. Concretely, DISTAR drafts block-level RVQ tokens with an AR language model and then performs parallel masked-diffusion infilling conditioned on the draft to complete the next block, yielding long-form synthesis with blockwise parallelism while mitigating classic AR exposure bias. The discrete code space affords explicit control at inference: DISTAR produces high-quality audio under both greedy and sample-based decoding using classifier-free guidance, supports trade-offs between robustness and diversity, and enables variable bit-rate and controllable computation via RVQ layer pruning at test time. Extensive experiments and ablations demonstrate that DISTAR surpasses state-of-the-art zero-shot TTS systems in robustness, naturalness, and speaker/style consistency, while maintaining rich output diversity. Audio samples are provided on https://anonymous.4open.science/w/DiSTAR_demo.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12210
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
Song, Yakun
Zhuang, Xiaobin
Chen, Jiawei
Niu, Zhikang
Yang, Guanrou
Du, Chenpeng
Jia, Dongya
Chen, Zhuo
Wang, Yuping
Wang, Yuxuan
Chen, Xie
Audio and Speech Processing
Computation and Language
Machine Learning
Recent attempts to interleave autoregressive (AR) sketchers with diffusion-based refiners over continuous speech representations have shown promise, but they remain brittle under distribution shift and offer limited levers for controllability. We introduce DISTAR, a zero-shot text-to-speech framework that operates entirely in a discrete residual vector quantization (RVQ) code space and tightly couples an AR language model with a masked diffusion model, without forced alignment or a duration predictor. Concretely, DISTAR drafts block-level RVQ tokens with an AR language model and then performs parallel masked-diffusion infilling conditioned on the draft to complete the next block, yielding long-form synthesis with blockwise parallelism while mitigating classic AR exposure bias. The discrete code space affords explicit control at inference: DISTAR produces high-quality audio under both greedy and sample-based decoding using classifier-free guidance, supports trade-offs between robustness and diversity, and enables variable bit-rate and controllable computation via RVQ layer pruning at test time. Extensive experiments and ablations demonstrate that DISTAR surpasses state-of-the-art zero-shot TTS systems in robustness, naturalness, and speaker/style consistency, while maintaining rich output diversity. Audio samples are provided on https://anonymous.4open.science/w/DiSTAR_demo.
title DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
topic Audio and Speech Processing
Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.12210