SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Wenxi, Wang, Xinsheng, Yan, Ruiqi, Chen, Yushen, Niu, Zhikang, Ma, Ziyang, Li, Xiquan, Liang, Yuzhe, Wen, Hanlin, Yin, Shunshun, Tao, Ming, Chen, Xie
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908707894329344
author Chen, Wenxi
Wang, Xinsheng
Yan, Ruiqi
Chen, Yushen
Niu, Zhikang
Ma, Ziyang
Li, Xiquan
Liang, Yuzhe
Wen, Hanlin
Yin, Shunshun
Tao, Ming
Chen, Xie
author_facet Chen, Wenxi
Wang, Xinsheng
Yan, Ruiqi
Chen, Yushen
Niu, Zhikang
Ma, Ziyang
Li, Xiquan
Liang, Yuzhe
Wen, Hanlin
Yin, Shunshun
Tao, Ming
Chen, Xie
contents Speech codecs that convert continuous speech signals into discrete tokens have become essential for speech language models. However, existing codecs struggle to balance high-quality reconstruction with semantically rich representations, limiting their effectiveness in both generative and understanding tasks. In this work, we propose SAC, a neural speech codec with semantic-acoustic dual-stream quantization. By disentangling semantic and acoustic modeling into two dedicated streams, SAC enables each to be optimized for its respective role. Comprehensive evaluations show that SAC achieves strong reconstruction performance across diverse bitrates under both clean and noisy conditions, with particularly high scores on UTMOS and WER, indicating superior naturalness and intelligibility. Moreover, SAC substantially surpasses prior codecs in semantic representation, approaching the level of continuous self-supervised embeddings. When used as a tokenizer for LLM-based text-to-speech, SAC enables a single-stage autoregressive (AR) TTS model that clearly outperforms state-of-the-art AR systems. Our disentanglement analysis further validates the effectiveness of the dual-stream design, offering new potential for controllable speech generation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16841
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
Chen, Wenxi
Wang, Xinsheng
Yan, Ruiqi
Chen, Yushen
Niu, Zhikang
Ma, Ziyang
Li, Xiquan
Liang, Yuzhe
Wen, Hanlin
Yin, Shunshun
Tao, Ming
Chen, Xie
Audio and Speech Processing
Sound
Speech codecs that convert continuous speech signals into discrete tokens have become essential for speech language models. However, existing codecs struggle to balance high-quality reconstruction with semantically rich representations, limiting their effectiveness in both generative and understanding tasks. In this work, we propose SAC, a neural speech codec with semantic-acoustic dual-stream quantization. By disentangling semantic and acoustic modeling into two dedicated streams, SAC enables each to be optimized for its respective role. Comprehensive evaluations show that SAC achieves strong reconstruction performance across diverse bitrates under both clean and noisy conditions, with particularly high scores on UTMOS and WER, indicating superior naturalness and intelligibility. Moreover, SAC substantially surpasses prior codecs in semantic representation, approaching the level of continuous self-supervised embeddings. When used as a tokenizer for LLM-based text-to-speech, SAC enables a single-stage autoregressive (AR) TTS model that clearly outperforms state-of-the-art AR systems. Our disentanglement analysis further validates the effectiveness of the dual-stream design, offering new potential for controllable speech generation.
title SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2510.16841