Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Haoran, Chen, Guanyu, Li, Bohan, Wang, Hankun, Guo, Yiwei, Li, Zhihan, Chen, Xie, Yu, Kai
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909626185809920
author Wang, Haoran
Chen, Guanyu
Li, Bohan
Wang, Hankun
Guo, Yiwei
Li, Zhihan
Chen, Xie
Yu, Kai
author_facet Wang, Haoran
Chen, Guanyu
Li, Bohan
Wang, Hankun
Guo, Yiwei
Li, Zhihan
Chen, Xie
Yu, Kai
contents Neural speech codecs excel in reconstructing clean speech signals; however, their efficacy in complex acoustic environments and downstream signal processing tasks remains underexplored. In this study, we introduce a novel benchmark named Environment-Resilient Speech Codec Benchmark (ERSB) to systematically evaluate whether neural speech codecs are environment-resilient. Specifically, we assess two key capabilities: (1) robust reconstruction, which measures the preservation of both speech and non-speech acoustic details, and (2) downstream task consistency, which ensures minimal deviation in downstream signal processing tasks when using reconstructed speech instead of the original. Our comprehensive experiments reveal that complex acoustic environments significantly degrade signal reconstruction and downstream task consistency. This work highlights the limitations of current speech codecs and raises a future direction that improves them for greater environmental resilience.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22515
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
Wang, Haoran
Chen, Guanyu
Li, Bohan
Wang, Hankun
Guo, Yiwei
Li, Zhihan
Chen, Xie
Yu, Kai
Sound
Audio and Speech Processing
Neural speech codecs excel in reconstructing clean speech signals; however, their efficacy in complex acoustic environments and downstream signal processing tasks remains underexplored. In this study, we introduce a novel benchmark named Environment-Resilient Speech Codec Benchmark (ERSB) to systematically evaluate whether neural speech codecs are environment-resilient. Specifically, we assess two key capabilities: (1) robust reconstruction, which measures the preservation of both speech and non-speech acoustic details, and (2) downstream task consistency, which ensures minimal deviation in downstream signal processing tasks when using reconstructed speech instead of the original. Our comprehensive experiments reveal that complex acoustic environments significantly degrade signal reconstruction and downstream task consistency. This work highlights the limitations of current speech codecs and raises a future direction that improves them for greater environmental resilience.
title Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.22515