A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918025745137664 |
|---|---|
| author | Xiang, Yang Huang, Canan Hu, Desheng Tian, Jingguang Hu, Xinhui Zhang, Chao |
| author_facet | Xiang, Yang Huang, Canan Hu, Desheng Tian, Jingguang Hu, Xinhui Zhang, Chao |
| contents | Most current speech enhancement (SE) methods recover clean speech from noisy inputs by directly estimating time-frequency masks or spectrums. However, these approaches often neglect the distinct attributes, such as semantic content and acoustic details, inherent in speech signals, which can hinder performance in downstream tasks. Moreover, their effectiveness tends to degrade in complex acoustic environments. To overcome these challenges, we propose a novel, semantic information-based, step-by-step factorized SE method using factorized codec and diffusion model. Unlike traditional SE methods, our hierarchical modeling of semantic and acoustic attributes enables more robust clean speech recovery, particularly in challenging acoustic scenarios. Moreover, this method offers further advantages for downstream TTS tasks. Experimental results demonstrate that our algorithm not only outperforms SOTA baselines in terms of speech quality but also enhances TTS performance in noisy environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_13843 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model Xiang, Yang Huang, Canan Hu, Desheng Tian, Jingguang Hu, Xinhui Zhang, Chao Audio and Speech Processing Sound Most current speech enhancement (SE) methods recover clean speech from noisy inputs by directly estimating time-frequency masks or spectrums. However, these approaches often neglect the distinct attributes, such as semantic content and acoustic details, inherent in speech signals, which can hinder performance in downstream tasks. Moreover, their effectiveness tends to degrade in complex acoustic environments. To overcome these challenges, we propose a novel, semantic information-based, step-by-step factorized SE method using factorized codec and diffusion model. Unlike traditional SE methods, our hierarchical modeling of semantic and acoustic attributes enables more robust clean speech recovery, particularly in challenging acoustic scenarios. Moreover, this method offers further advantages for downstream TTS tasks. Experimental results demonstrate that our algorithm not only outperforms SOTA baselines in terms of speech quality but also enhances TTS performance in noisy environments. |
| title | A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2505.13843 |