Optimizing Neural Speech Codec for Low-Bitrate Compression via Multi-Scale Encoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Peiji, Wang, Fengping, Zhong, Yicheng, Wei, Huawei, Wang, Zhisheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909357445218304
author Yang, Peiji
Wang, Fengping
Zhong, Yicheng
Wei, Huawei
Wang, Zhisheng
author_facet Yang, Peiji
Wang, Fengping
Zhong, Yicheng
Wei, Huawei
Wang, Zhisheng
contents Neural speech codecs have demonstrated their ability to compress high-quality speech and audio by converting them into discrete token representations. Most existing methods utilize Residual Vector Quantization (RVQ) to encode speech into multiple layers of discrete codes with uniform time scales. However, this strategy overlooks the differences in information density across various speech features, leading to redundant encoding of sparse information, which limits the performance of these methods at low bitrate. This paper proposes MsCodec, a novel multi-scale neural speech codec that encodes speech into multiple layers of discrete codes, each corresponding to a different time scale. This encourages the model to decouple speech features according to their diverse information densities, consequently enhancing the performance of speech compression. Furthermore, we incorporate mutual information loss to augment the diversity among speech codes across different layers. Experimental results indicate that our proposed method significantly improves codec performance at low bitrate.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15749
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Optimizing Neural Speech Codec for Low-Bitrate Compression via Multi-Scale Encoding
Yang, Peiji
Wang, Fengping
Zhong, Yicheng
Wei, Huawei
Wang, Zhisheng
Sound
Audio and Speech Processing
Neural speech codecs have demonstrated their ability to compress high-quality speech and audio by converting them into discrete token representations. Most existing methods utilize Residual Vector Quantization (RVQ) to encode speech into multiple layers of discrete codes with uniform time scales. However, this strategy overlooks the differences in information density across various speech features, leading to redundant encoding of sparse information, which limits the performance of these methods at low bitrate. This paper proposes MsCodec, a novel multi-scale neural speech codec that encodes speech into multiple layers of discrete codes, each corresponding to a different time scale. This encourages the model to decouple speech features according to their diverse information densities, consequently enhancing the performance of speech compression. Furthermore, we incorporate mutual information loss to augment the diversity among speech codes across different layers. Experimental results indicate that our proposed method significantly improves codec performance at low bitrate.
title Optimizing Neural Speech Codec for Low-Bitrate Compression via Multi-Scale Encoding
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.15749