A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Alexander H., Wang, Qirui, Gong, Yuan, Glass, James
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913567271288832
author Liu, Alexander H.
Wang, Qirui
Gong, Yuan
Glass, James
author_facet Liu, Alexander H.
Wang, Qirui
Gong, Yuan
Glass, James
contents Neural Audio Codecs, initially designed as a compression technique, have gained more attention recently for speech generation. Codec models represent each audio frame as a sequence of tokens, i.e., discrete embeddings. The discrete and low-frequency nature of neural codecs introduced a new way to generate speech with token-based models. As these tokens encode information at various levels of granularity, from coarse to fine, most existing works focus on how to better generate the coarse tokens. In this paper, we focus on an equally important but often overlooked question: How can we better resynthesize the waveform from coarse tokens? We point out that both the choice of learning target and resynthesis approach have a dramatic impact on the generated audio quality. Specifically, we study two different strategies based on token prediction and regression, and introduce a new method based on Schrödinger Bridge. We examine how different design choices affect machine and human perception.
format Preprint
id arxiv_https___arxiv_org_abs_2410_22448
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
Liu, Alexander H.
Wang, Qirui
Gong, Yuan
Glass, James
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
Neural Audio Codecs, initially designed as a compression technique, have gained more attention recently for speech generation. Codec models represent each audio frame as a sequence of tokens, i.e., discrete embeddings. The discrete and low-frequency nature of neural codecs introduced a new way to generate speech with token-based models. As these tokens encode information at various levels of granularity, from coarse to fine, most existing works focus on how to better generate the coarse tokens. In this paper, we focus on an equally important but often overlooked question: How can we better resynthesize the waveform from coarse tokens? We point out that both the choice of learning target and resynthesis approach have a dramatic impact on the generated audio quality. Specifically, we study two different strategies based on token prediction and regression, and introduce a new method based on Schrödinger Bridge. We examine how different design choices affect machine and human perception.
title A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2410.22448