Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Tan Dat, Kim, Ji-Hoon, Choi, Jeongsoo, Choi, Shukjae, Park, Jinseok, Lee, Younglo, Chung, Joon Son
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910654673190912
author Nguyen, Tan Dat
Kim, Ji-Hoon
Choi, Jeongsoo
Choi, Shukjae
Park, Jinseok
Lee, Younglo
Chung, Joon Son
author_facet Nguyen, Tan Dat
Kim, Ji-Hoon
Choi, Jeongsoo
Choi, Shukjae
Park, Jinseok
Lee, Younglo
Chung, Joon Son
contents The goal of this paper is to accelerate codec-based speech synthesis systems with minimum sacrifice to speech quality. We propose an enhanced inference method that allows for flexible trade-offs between speed and quality during inference without requiring additional training. Our core idea is to predict multiple tokens per inference step of the AR module using multiple prediction heads, resulting in a linear reduction in synthesis time as the number of heads increases. Furthermore, we introduce a novel speculative decoding technique that utilises a Viterbi-based algorithm to select the optimal sequence of generated tokens at each decoding step. In our experiments, we demonstrate that the time required to predict each token is reduced by a factor of 4 to 5 compared to baseline models, with minimal quality trade-off or even improvement in terms of speech intelligibility. Audio samples are available at: multpletokensprediction.github.io/multipletokensprediction.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13839
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
Nguyen, Tan Dat
Kim, Ji-Hoon
Choi, Jeongsoo
Choi, Shukjae
Park, Jinseok
Lee, Younglo
Chung, Joon Son
Sound
Artificial Intelligence
Audio and Speech Processing
The goal of this paper is to accelerate codec-based speech synthesis systems with minimum sacrifice to speech quality. We propose an enhanced inference method that allows for flexible trade-offs between speed and quality during inference without requiring additional training. Our core idea is to predict multiple tokens per inference step of the AR module using multiple prediction heads, resulting in a linear reduction in synthesis time as the number of heads increases. Furthermore, we introduce a novel speculative decoding technique that utilises a Viterbi-based algorithm to select the optimal sequence of generated tokens at each decoding step. In our experiments, we demonstrate that the time required to predict each token is reduced by a factor of 4 to 5 compared to baseline models, with minimal quality trade-off or even improvement in terms of speech intelligibility. Audio samples are available at: multpletokensprediction.github.io/multipletokensprediction.github.io/.
title Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2410.13839