Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Bohan, Li, Zhihan, Wang, Haoran, Zhang, Hanglei, Guo, Yiwei, Wang, Hankun, Chen, Xie, Yu, Kai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911026038964224
author Li, Bohan
Li, Zhihan
Wang, Haoran
Zhang, Hanglei
Guo, Yiwei
Wang, Hankun
Chen, Xie
Yu, Kai
author_facet Li, Bohan
Li, Zhihan
Wang, Haoran
Zhang, Hanglei
Guo, Yiwei
Wang, Hankun
Chen, Xie
Yu, Kai
contents Recently, autoregressive (AR) language models have emerged as a dominant approach in speech synthesis, offering expressive generation and scalable training. However, conventional AR speech synthesis models relying on the next-token prediction paradigm often encounter significant challenges when handling long speech sequences. These models often struggle to construct stable frame-to-frame attention, leading to increased latency and degraded synthesis quality, thereby limiting their feasibility for real-time applications. To address these limitations, we introduce a novel dynamic chunk-wise autoregressive synthesis framework, termed DCAR, designed to enhance both efficiency and intelligibility robustness in AR speech generation. DCAR introduces a chunk-to-frame attention mechanism through training with multi-token prediction, enabling dynamic chunk prediction in variable speech contexts using a lightweight module trained on-policy. DCAR dynamically adjusts the token prediction span, significantly reducing the sequence length dependency while obtaining high synthesis quality. Comprehensive empirical evaluations demonstrate that DCAR substantially outperforms traditional next-token prediction models, achieving up to 72.27% intelligibility improvement and 2.61x inference speedup simultaneously on the test set. Furthermore, we conduct comprehensive analysis to support it as a versatile foundation for next-generation speech synthesis systems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22023
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
Li, Bohan
Li, Zhihan
Wang, Haoran
Zhang, Hanglei
Guo, Yiwei
Wang, Hankun
Chen, Xie
Yu, Kai
Sound
Computation and Language
Audio and Speech Processing
Recently, autoregressive (AR) language models have emerged as a dominant approach in speech synthesis, offering expressive generation and scalable training. However, conventional AR speech synthesis models relying on the next-token prediction paradigm often encounter significant challenges when handling long speech sequences. These models often struggle to construct stable frame-to-frame attention, leading to increased latency and degraded synthesis quality, thereby limiting their feasibility for real-time applications. To address these limitations, we introduce a novel dynamic chunk-wise autoregressive synthesis framework, termed DCAR, designed to enhance both efficiency and intelligibility robustness in AR speech generation. DCAR introduces a chunk-to-frame attention mechanism through training with multi-token prediction, enabling dynamic chunk prediction in variable speech contexts using a lightweight module trained on-policy. DCAR dynamically adjusts the token prediction span, significantly reducing the sequence length dependency while obtaining high synthesis quality. Comprehensive empirical evaluations demonstrate that DCAR substantially outperforms traditional next-token prediction models, achieving up to 72.27% intelligibility improvement and 2.61x inference speedup simultaneously on the test set. Furthermore, we conduct comprehensive analysis to support it as a versatile foundation for next-generation speech synthesis systems.
title Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2506.22023