SaD: A Scenario-Aware Discriminator for Speech Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Xihao, Liu, Siqi, Chen, Yan, Zhou, Hang, Liu, Chang, Chen, Hanting, Hu, Jie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912577483702272
author Yuan, Xihao
Liu, Siqi
Chen, Yan
Zhou, Hang
Liu, Chang
Chen, Hanting
Hu, Jie
author_facet Yuan, Xihao
Liu, Siqi
Chen, Yan
Zhou, Hang
Liu, Chang
Chen, Hanting
Hu, Jie
contents Generative adversarial network-based models have shown remarkable performance in the field of speech enhancement. However, the current optimization strategies for these models predominantly focus on refining the architecture of the generator or enhancing the quality evaluation metrics of the discriminator. This approach often overlooks the rich contextual information inherent in diverse scenarios. In this paper, we propose a scenario-aware discriminator that captures scene-specific features and performs frequency-domain division, thereby enabling a more accurate quality assessment of the enhanced speech generated by the generator. We conducted comprehensive experiments on three representative models using two publicly available datasets. The results demonstrate that our method can effectively adapt to various generator architectures without altering their structure, thereby unlocking further performance gains in speech enhancement across different scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00405
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SaD: A Scenario-Aware Discriminator for Speech Enhancement
Yuan, Xihao
Liu, Siqi
Chen, Yan
Zhou, Hang
Liu, Chang
Chen, Hanting
Hu, Jie
Sound
Audio and Speech Processing
Generative adversarial network-based models have shown remarkable performance in the field of speech enhancement. However, the current optimization strategies for these models predominantly focus on refining the architecture of the generator or enhancing the quality evaluation metrics of the discriminator. This approach often overlooks the rich contextual information inherent in diverse scenarios. In this paper, we propose a scenario-aware discriminator that captures scene-specific features and performs frequency-domain division, thereby enabling a more accurate quality assessment of the enhanced speech generated by the generator. We conducted comprehensive experiments on three representative models using two publicly available datasets. The results demonstrate that our method can effectively adapt to various generator architectures without altering their structure, thereby unlocking further performance gains in speech enhancement across different scenarios.
title SaD: A Scenario-Aware Discriminator for Speech Enhancement
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.00405