Efficient Evaluation of Quantization-Effects in Neural Codecs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mack, Wolfgang, Mustafa, Ahmed, Łaganowski, Rafał, Hijazy, Samer
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915142135971840
author Mack, Wolfgang
Mustafa, Ahmed
Łaganowski, Rafał
Hijazy, Samer
author_facet Mack, Wolfgang
Mustafa, Ahmed
Łaganowski, Rafał
Hijazy, Samer
contents Neural codecs, comprising an encoder, quantizer, and decoder, enable signal transmission at exceptionally low bitrates. Training these systems requires techniques like the straight-through estimator, soft-to-hard annealing, or statistical quantizer emulation to allow a non-zero gradient across the quantizer. Evaluating the effect of quantization in neural codecs, like the influence of gradient passing techniques on the whole system, is often costly and time-consuming due to training demands and the lack of affordable and reliable metrics. This paper proposes an efficient evaluation framework for neural codecs using simulated data with a defined number of bits and low-complexity neural encoders/decoders to emulate the non-linear behavior in larger networks. Our system is highly efficient in terms of training time and computational and hardware requirements, allowing us to uncover distinct behaviors in neural codecs. We propose a modification to stabilize training with the straight-through estimator based on our findings. We validate our findings against an internal neural audio codec and against the state-of-the-art descript-audio-codec.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04770
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Evaluation of Quantization-Effects in Neural Codecs
Mack, Wolfgang
Mustafa, Ahmed
Łaganowski, Rafał
Hijazy, Samer
Audio and Speech Processing
Machine Learning
Neural codecs, comprising an encoder, quantizer, and decoder, enable signal transmission at exceptionally low bitrates. Training these systems requires techniques like the straight-through estimator, soft-to-hard annealing, or statistical quantizer emulation to allow a non-zero gradient across the quantizer. Evaluating the effect of quantization in neural codecs, like the influence of gradient passing techniques on the whole system, is often costly and time-consuming due to training demands and the lack of affordable and reliable metrics. This paper proposes an efficient evaluation framework for neural codecs using simulated data with a defined number of bits and low-complexity neural encoders/decoders to emulate the non-linear behavior in larger networks. Our system is highly efficient in terms of training time and computational and hardware requirements, allowing us to uncover distinct behaviors in neural codecs. We propose a modification to stabilize training with the straight-through estimator based on our findings. We validate our findings against an internal neural audio codec and against the state-of-the-art descript-audio-codec.
title Efficient Evaluation of Quantization-Effects in Neural Codecs
topic Audio and Speech Processing
Machine Learning
url https://arxiv.org/abs/2502.04770