HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xue, Rongkun, Niu, Yazhe, Hu, Shuai, Yin, Zixin, Yao, Yongqiang, Yang, Jing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913959628505088
author Xue, Rongkun
Niu, Yazhe
Hu, Shuai
Yin, Zixin
Yao, Yongqiang
Yang, Jing
author_facet Xue, Rongkun
Niu, Yazhe
Hu, Shuai
Yin, Zixin
Yao, Yongqiang
Yang, Jing
contents Discrete speech tokenization is a fundamental component in speech codecs. However, in large-scale speech-to-speech systems, the complexity of parallel streams from multiple quantizers and the computational cost of high-time-dimensional codecs pose significant challenges. In this paper, we introduce HH-Codec, a neural codec that achieves extreme compression at 24 tokens per second for 24 kHz audio while relying on single-quantizer inference. Our approach involves a carefully designed Vector Quantization space for Spoken Language Modeling, optimizing compression efficiency while minimizing information loss. Building on this, we propose an asymmetric encoder-decoder architecture (Audio-VQ-Mel-Audio) that leverages dual supervision and progressive training to enhance reconstruction stability and fidelity. HH-Codec achieves state-of-the-art performance in speech reconstruction with an ultra-low bandwidth of 0.3 kbps. We further evaluate its effectiveness in codebook utilization and generative model adaptation, with extensive ablations validating the necessity of each module. HH-Codec is available at https://github.com/opendilab/HH-Codec.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18897
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling
Xue, Rongkun
Niu, Yazhe
Hu, Shuai
Yin, Zixin
Yao, Yongqiang
Yang, Jing
Sound
Artificial Intelligence
Audio and Speech Processing
Discrete speech tokenization is a fundamental component in speech codecs. However, in large-scale speech-to-speech systems, the complexity of parallel streams from multiple quantizers and the computational cost of high-time-dimensional codecs pose significant challenges. In this paper, we introduce HH-Codec, a neural codec that achieves extreme compression at 24 tokens per second for 24 kHz audio while relying on single-quantizer inference. Our approach involves a carefully designed Vector Quantization space for Spoken Language Modeling, optimizing compression efficiency while minimizing information loss. Building on this, we propose an asymmetric encoder-decoder architecture (Audio-VQ-Mel-Audio) that leverages dual supervision and progressive training to enhance reconstruction stability and fidelity. HH-Codec achieves state-of-the-art performance in speech reconstruction with an ultra-low bandwidth of 0.3 kbps. We further evaluate its effectiveness in codebook utilization and generative model adaptation, with extensive ablations validating the necessity of each module. HH-Codec is available at https://github.com/opendilab/HH-Codec.
title HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2507.18897