QINCODEC: Neural Audio Compression with Implicit Neural Codebooks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lahrichi, Zineb, Hadjeres, Gaëtan, Richard, Gael, Peeters, Geoffroy
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909551856451584
author Lahrichi, Zineb
Hadjeres, Gaëtan
Richard, Gael
Peeters, Geoffroy
author_facet Lahrichi, Zineb
Hadjeres, Gaëtan
Richard, Gael
Peeters, Geoffroy
contents Neural audio codecs, neural networks which compress a waveform into discrete tokens, play a crucial role in the recent development of audio generative models. State-of-the-art codecs rely on the end-to-end training of an autoencoder and a quantization bottleneck. However, this approach restricts the choice of the quantization methods as it requires to define how gradients propagate through the quantizer and how to update the quantization parameters online. In this work, we revisit the common practice of joint training and propose to quantize the latent representations of a pre-trained autoencoder offline, followed by an optional finetuning of the decoder to mitigate degradation from quantization. This strategy allows to consider any off-the-shelf quantizer, especially state-of-the-art trainable quantizers with implicit neural codebooks such as QINCO2. We demonstrate that with the latter, our proposed codec termed QINCODEC, is competitive with baseline codecs while being notably simpler to train. Finally, our approach provides a general framework that amortizes the cost of autoencoder pretraining, and enables more flexible codec design.
format Preprint
id arxiv_https___arxiv_org_abs_2503_19597
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QINCODEC: Neural Audio Compression with Implicit Neural Codebooks
Lahrichi, Zineb
Hadjeres, Gaëtan
Richard, Gael
Peeters, Geoffroy
Sound
Signal Processing
Neural audio codecs, neural networks which compress a waveform into discrete tokens, play a crucial role in the recent development of audio generative models. State-of-the-art codecs rely on the end-to-end training of an autoencoder and a quantization bottleneck. However, this approach restricts the choice of the quantization methods as it requires to define how gradients propagate through the quantizer and how to update the quantization parameters online. In this work, we revisit the common practice of joint training and propose to quantize the latent representations of a pre-trained autoencoder offline, followed by an optional finetuning of the decoder to mitigate degradation from quantization. This strategy allows to consider any off-the-shelf quantizer, especially state-of-the-art trainable quantizers with implicit neural codebooks such as QINCO2. We demonstrate that with the latter, our proposed codec termed QINCODEC, is competitive with baseline codecs while being notably simpler to train. Finally, our approach provides a general framework that amortizes the cost of autoencoder pretraining, and enables more flexible codec design.
title QINCODEC: Neural Audio Compression with Implicit Neural Codebooks
topic Sound
Signal Processing
url https://arxiv.org/abs/2503.19597