DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Jiaqi, Lin, Xiaolong, Li, Zhekai, Huang, Shixi, Wang, Yuancheng, Wang, Chaoren, Zhan, Zhenpeng, Wu, Zhizheng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912619713003520
author Li, Jiaqi
Lin, Xiaolong
Li, Zhekai
Huang, Shixi
Wang, Yuancheng
Wang, Chaoren
Zhan, Zhenpeng
Wu, Zhizheng
author_facet Li, Jiaqi
Lin, Xiaolong
Li, Zhekai
Huang, Shixi
Wang, Yuancheng
Wang, Chaoren
Zhan, Zhenpeng
Wu, Zhizheng
contents Neural audio codecs form the foundational building blocks for language model (LM)-based speech generation. Typically, there is a trade-off between frame rate and audio quality. This study introduces a low-frame-rate, semantically enhanced codec model. Existing approaches distill semantically rich self-supervised (SSL) representations into the first-layer codec tokens. This work proposes DualCodec, a dual-stream encoding approach that integrates SSL and waveform representations within an end-to-end codec framework. In this setting, DualCodec enhances the semantic information in the first-layer codec and enables the codec system to maintain high audio quality while operating at a low frame rate. Note that a low-frame-rate codec improves the efficiency of speech generation. Experimental results on audio codec and speech generation tasks confirm the effectiveness of the proposed DualCodec compared to state-of-the-art codec systems, such as Mimi Codec, SpeechTokenizer, DAC, and Encodec. Demos are available at: https://dualcodec.github.io, code is available at: https://github.com/jiaqili3/DualCodec
format Preprint
id arxiv_https___arxiv_org_abs_2505_13000
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
Li, Jiaqi
Lin, Xiaolong
Li, Zhekai
Huang, Shixi
Wang, Yuancheng
Wang, Chaoren
Zhan, Zhenpeng
Wu, Zhizheng
Sound
Audio and Speech Processing
Neural audio codecs form the foundational building blocks for language model (LM)-based speech generation. Typically, there is a trade-off between frame rate and audio quality. This study introduces a low-frame-rate, semantically enhanced codec model. Existing approaches distill semantically rich self-supervised (SSL) representations into the first-layer codec tokens. This work proposes DualCodec, a dual-stream encoding approach that integrates SSL and waveform representations within an end-to-end codec framework. In this setting, DualCodec enhances the semantic information in the first-layer codec and enables the codec system to maintain high audio quality while operating at a low frame rate. Note that a low-frame-rate codec improves the efficiency of speech generation. Experimental results on audio codec and speech generation tasks confirm the effectiveness of the proposed DualCodec compared to state-of-the-art codec systems, such as Mimi Codec, SpeechTokenizer, DAC, and Encodec. Demos are available at: https://dualcodec.github.io, code is available at: https://github.com/jiaqili3/DualCodec
title DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.13000