NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Niu, Zhikang, Chen, Sanyuan, Zhou, Long, Ma, Ziyang, Chen, Xie, Liu, Shujie
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913508270014464
author Niu, Zhikang
Chen, Sanyuan
Zhou, Long
Ma, Ziyang
Chen, Xie
Liu, Shujie
author_facet Niu, Zhikang
Chen, Sanyuan
Zhou, Long
Ma, Ziyang
Chen, Xie
Liu, Shujie
contents Built upon vector quantization (VQ), discrete audio codec models have achieved great success in audio compression and auto-regressive audio generation. However, existing models face substantial challenges in perceptual quality and signal distortion, especially when operating in extremely low bandwidth, rooted in the sensitivity of the VQ codebook to noise. This degradation poses significant challenges for several downstream tasks, such as codec-based speech synthesis. To address this issue, we propose a novel VQ method, Normal Distribution-based Vector Quantization (NDVQ), by introducing an explicit margin between the VQ codes via learning a variance. Specifically, our approach involves mapping the waveform to a latent space and quantizing it by selecting the most likely normal distribution, with each codebook entry representing a unique normal distribution defined by its mean and variance. Using these distribution-based VQ codec codes, a decoder reconstructs the input waveform. NDVQ is trained with additional distribution-related losses, alongside reconstruction and discrimination losses. Experiments demonstrate that NDVQ outperforms existing audio compression baselines, such as EnCodec, in terms of audio quality and zero-shot TTS, particularly in very low bandwidth scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2409_12717
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
Niu, Zhikang
Chen, Sanyuan
Zhou, Long
Ma, Ziyang
Chen, Xie
Liu, Shujie
Audio and Speech Processing
Sound
Built upon vector quantization (VQ), discrete audio codec models have achieved great success in audio compression and auto-regressive audio generation. However, existing models face substantial challenges in perceptual quality and signal distortion, especially when operating in extremely low bandwidth, rooted in the sensitivity of the VQ codebook to noise. This degradation poses significant challenges for several downstream tasks, such as codec-based speech synthesis. To address this issue, we propose a novel VQ method, Normal Distribution-based Vector Quantization (NDVQ), by introducing an explicit margin between the VQ codes via learning a variance. Specifically, our approach involves mapping the waveform to a latent space and quantizing it by selecting the most likely normal distribution, with each codebook entry representing a unique normal distribution defined by its mean and variance. Using these distribution-based VQ codec codes, a decoder reconstructs the input waveform. NDVQ is trained with additional distribution-related losses, alongside reconstruction and discrimination losses. Experiments demonstrate that NDVQ outperforms existing audio compression baselines, such as EnCodec, in terms of audio quality and zero-shot TTS, particularly in very low bandwidth scenarios.
title NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2409.12717