SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Sirui, Wang, Shuai, Liu, Zhijun, Jiang, Zhongjie, Wang, Yannan, Li, Haizhou
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916796049653760
author Li, Sirui
Wang, Shuai
Liu, Zhijun
Jiang, Zhongjie
Wang, Yannan
Li, Haizhou
author_facet Li, Sirui
Wang, Shuai
Liu, Zhijun
Jiang, Zhongjie
Wang, Yannan
Li, Haizhou
contents Speech pre-processing techniques such as denoising, de-reverberation, and separation, are commonly employed as front-ends for various downstream speech processing tasks. However, these methods can sometimes be inadequate, resulting in residual noise or the introduction of new artifacts. Such deficiencies are typically not captured by metrics like SI-SNR but are noticeable to human listeners. To address this, we introduce SpeechRefiner, a post-processing tool that utilizes Conditional Flow Matching (CFM) to improve the perceptual quality of speech. In this study, we benchmark SpeechRefiner against recent task-specific refinement methods and evaluate its performance within our internal processing pipeline, which integrates multiple front-end algorithms. Experiments show that SpeechRefiner exhibits strong generalization across diverse impairment sources, significantly enhancing speech perceptual quality. Audio demos can be found at https://speechrefiner.github.io/SpeechRefiner/.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13709
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
Li, Sirui
Wang, Shuai
Liu, Zhijun
Jiang, Zhongjie
Wang, Yannan
Li, Haizhou
Audio and Speech Processing
Sound
Speech pre-processing techniques such as denoising, de-reverberation, and separation, are commonly employed as front-ends for various downstream speech processing tasks. However, these methods can sometimes be inadequate, resulting in residual noise or the introduction of new artifacts. Such deficiencies are typically not captured by metrics like SI-SNR but are noticeable to human listeners. To address this, we introduce SpeechRefiner, a post-processing tool that utilizes Conditional Flow Matching (CFM) to improve the perceptual quality of speech. In this study, we benchmark SpeechRefiner against recent task-specific refinement methods and evaluate its performance within our internal processing pipeline, which integrates multiple front-end algorithms. Experiments show that SpeechRefiner exhibits strong generalization across diverse impairment sources, significantly enhancing speech perceptual quality. Audio demos can be found at https://speechrefiner.github.io/SpeechRefiner/.
title SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2506.13709