Interspeech 2025 URGENT Speech Enhancement Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Saijo, Kohei, Zhang, Wangyou, Cornell, Samuele, Scheibler, Robin, Li, Chenda, Ni, Zhaoheng, Kumar, Anurag, Sach, Marvin, Fu, Yihui, Wang, Wei, Fingscheidt, Tim, Watanabe, Shinji
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918042657619968
author Saijo, Kohei
Zhang, Wangyou
Cornell, Samuele
Scheibler, Robin
Li, Chenda
Ni, Zhaoheng
Kumar, Anurag
Sach, Marvin
Fu, Yihui
Wang, Wei
Fingscheidt, Tim
Watanabe, Shinji
author_facet Saijo, Kohei
Zhang, Wangyou
Cornell, Samuele
Scheibler, Robin
Li, Chenda
Ni, Zhaoheng
Kumar, Anurag
Sach, Marvin
Fu, Yihui
Wang, Wei
Fingscheidt, Tim
Watanabe, Shinji
contents There has been a growing effort to develop universal speech enhancement (SE) to handle inputs with various speech distortions and recording conditions. The URGENT Challenge series aims to foster such universal SE by embracing a broad range of distortion types, increasing data diversity, and incorporating extensive evaluation metrics. This work introduces the Interspeech 2025 URGENT Challenge, the second edition of the series, to explore several aspects that have received limited attention so far: language dependency, universality for more distortion types, data scalability, and the effectiveness of using noisy training data. We received 32 submissions, where the best system uses a discriminative model, while most other competitive ones are hybrid methods. Analysis reveals some key findings: (i) some generative or hybrid approaches are preferred in subjective evaluations over the top discriminative model, and (ii) purely generative SE models can exhibit language dependency.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23212
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Interspeech 2025 URGENT Speech Enhancement Challenge
Saijo, Kohei
Zhang, Wangyou
Cornell, Samuele
Scheibler, Robin
Li, Chenda
Ni, Zhaoheng
Kumar, Anurag
Sach, Marvin
Fu, Yihui
Wang, Wei
Fingscheidt, Tim
Watanabe, Shinji
Audio and Speech Processing
There has been a growing effort to develop universal speech enhancement (SE) to handle inputs with various speech distortions and recording conditions. The URGENT Challenge series aims to foster such universal SE by embracing a broad range of distortion types, increasing data diversity, and incorporating extensive evaluation metrics. This work introduces the Interspeech 2025 URGENT Challenge, the second edition of the series, to explore several aspects that have received limited attention so far: language dependency, universality for more distortion types, data scalability, and the effectiveness of using noisy training data. We received 32 submissions, where the best system uses a discriminative model, while most other competitive ones are hybrid methods. Analysis reveals some key findings: (i) some generative or hybrid approaches are preferred in subjective evaluations over the top discriminative model, and (ii) purely generative SE models can exhibit language dependency.
title Interspeech 2025 URGENT Speech Enhancement Challenge
topic Audio and Speech Processing
url https://arxiv.org/abs/2505.23212