A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiang, Yang, Huang, Canan, Hu, Desheng, Tian, Jingguang, Hu, Xinhui, Zhang, Chao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918025745137664
author Xiang, Yang
Huang, Canan
Hu, Desheng
Tian, Jingguang
Hu, Xinhui
Zhang, Chao
author_facet Xiang, Yang
Huang, Canan
Hu, Desheng
Tian, Jingguang
Hu, Xinhui
Zhang, Chao
contents Most current speech enhancement (SE) methods recover clean speech from noisy inputs by directly estimating time-frequency masks or spectrums. However, these approaches often neglect the distinct attributes, such as semantic content and acoustic details, inherent in speech signals, which can hinder performance in downstream tasks. Moreover, their effectiveness tends to degrade in complex acoustic environments. To overcome these challenges, we propose a novel, semantic information-based, step-by-step factorized SE method using factorized codec and diffusion model. Unlike traditional SE methods, our hierarchical modeling of semantic and acoustic attributes enables more robust clean speech recovery, particularly in challenging acoustic scenarios. Moreover, this method offers further advantages for downstream TTS tasks. Experimental results demonstrate that our algorithm not only outperforms SOTA baselines in terms of speech quality but also enhances TTS performance in noisy environments.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13843
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model
Xiang, Yang
Huang, Canan
Hu, Desheng
Tian, Jingguang
Hu, Xinhui
Zhang, Chao
Audio and Speech Processing
Sound
Most current speech enhancement (SE) methods recover clean speech from noisy inputs by directly estimating time-frequency masks or spectrums. However, these approaches often neglect the distinct attributes, such as semantic content and acoustic details, inherent in speech signals, which can hinder performance in downstream tasks. Moreover, their effectiveness tends to degrade in complex acoustic environments. To overcome these challenges, we propose a novel, semantic information-based, step-by-step factorized SE method using factorized codec and diffusion model. Unlike traditional SE methods, our hierarchical modeling of semantic and acoustic attributes enables more robust clean speech recovery, particularly in challenging acoustic scenarios. Moreover, this method offers further advantages for downstream TTS tasks. Experimental results demonstrate that our algorithm not only outperforms SOTA baselines in terms of speech quality but also enhances TTS performance in noisy environments.
title A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2505.13843