Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Wei, Zhang, Wangyou, Li, Chenda, Shi, Jiatong, Watanabe, Shinji, Qian, Yanmin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916911674032128
author Wang, Wei
Zhang, Wangyou
Li, Chenda
Shi, Jiatong
Watanabe, Shinji
Qian, Yanmin
author_facet Wang, Wei
Zhang, Wangyou
Li, Chenda
Shi, Jiatong
Watanabe, Shinji
Qian, Yanmin
contents Speech quality assessment (SQA) aims to predict the perceived quality of speech signals under a wide range of distortions. It is inherently connected to speech enhancement (SE), which seeks to improve speech quality by removing unwanted signal components. While SQA models are widely used to evaluate SE performance, their potential to guide SE training remains underexplored. In this work, we investigate a training framework that leverages a SQA model, trained to predict multiple evaluation metrics from a public SE leaderboard, as a supervisory signal for SE. This approach addresses a key limitation of conventional SE objectives, such as SI-SNR, which often fail to align with perceptual quality and generalize poorly across evaluation metrics. Moreover, it enables training on real-world data where clean references are unavailable. Experiments on both simulated and real-world test sets show that SQA-guided training consistently improves performance across a range of quality metrics. Code and checkpoints are available at https://github.com/urgent-challenge/urgent2026_challenge_track2
format Preprint
id arxiv_https___arxiv_org_abs_2506_12260
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
Wang, Wei
Zhang, Wangyou
Li, Chenda
Shi, Jiatong
Watanabe, Shinji
Qian, Yanmin
Sound
Audio and Speech Processing
Speech quality assessment (SQA) aims to predict the perceived quality of speech signals under a wide range of distortions. It is inherently connected to speech enhancement (SE), which seeks to improve speech quality by removing unwanted signal components. While SQA models are widely used to evaluate SE performance, their potential to guide SE training remains underexplored. In this work, we investigate a training framework that leverages a SQA model, trained to predict multiple evaluation metrics from a public SE leaderboard, as a supervisory signal for SE. This approach addresses a key limitation of conventional SE objectives, such as SI-SNR, which often fail to align with perceptual quality and generalize poorly across evaluation metrics. Moreover, it enables training on real-world data where clean references are unavailable. Experiments on both simulated and real-world test sets show that SQA-guided training consistently improves performance across a range of quality metrics. Code and checkpoints are available at https://github.com/urgent-challenge/urgent2026_challenge_track2
title Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.12260