Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ren, Yanzhou, Harada, Noboru, Takeuchi, Daiki, Chen, Siyu, Liu, Wei, Zhang, Xiao, Zhang, Liyuan, Moriya, Takehiro, Makino, Shoji
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917305510789120
author Ren, Yanzhou
Harada, Noboru
Takeuchi, Daiki
Chen, Siyu
Liu, Wei
Zhang, Xiao
Zhang, Liyuan
Moriya, Takehiro
Makino, Shoji
author_facet Ren, Yanzhou
Harada, Noboru
Takeuchi, Daiki
Chen, Siyu
Liu, Wei
Zhang, Xiao
Zhang, Liyuan
Moriya, Takehiro
Makino, Shoji
contents Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensuring accurate semantic modeling while maintaining high-fidelity reconstruction under ultra-low bitrate constraints remains challenging. We propose an entropy-guided group residual vector quantization (EG-GRVQ) for an ultra-low bitrate neural speech codec, which retains a semantic branch for linguistic information and incorporates an entropy-guided grouping strategy in the acoustic branch. Assuming that channel activations follow approximately Gaussian statistics, the variance of each channel can serve as a principled proxy for its information content. Based on this assumption, we partition the encoder output such that each group carries an equal share of the total information. This balanced allocation improves codebook efficiency and reduces redundancy. Trained on LibriTTS and VCTK, our model shows improvements in perceptual quality and intelligibility metrics under ultra-low bitrate conditions, with a focus on codec-level fidelity for communication-oriented scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01476
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
Ren, Yanzhou
Harada, Noboru
Takeuchi, Daiki
Chen, Siyu
Liu, Wei
Zhang, Xiao
Zhang, Liyuan
Moriya, Takehiro
Makino, Shoji
Audio and Speech Processing
Signal Processing
Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensuring accurate semantic modeling while maintaining high-fidelity reconstruction under ultra-low bitrate constraints remains challenging. We propose an entropy-guided group residual vector quantization (EG-GRVQ) for an ultra-low bitrate neural speech codec, which retains a semantic branch for linguistic information and incorporates an entropy-guided grouping strategy in the acoustic branch. Assuming that channel activations follow approximately Gaussian statistics, the variance of each channel can serve as a principled proxy for its information content. Based on this assumption, we partition the encoder output such that each group carries an equal share of the total information. This balanced allocation improves codebook efficiency and reduces redundancy. Trained on LibriTTS and VCTK, our model shows improvements in perceptual quality and intelligibility metrics under ultra-low bitrate conditions, with a focus on codec-level fidelity for communication-oriented scenarios.
title Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
topic Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2603.01476