URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Wangyou, Scheibler, Robin, Saijo, Kohei, Cornell, Samuele, Li, Chenda, Ni, Zhaoheng, Kumar, Anurag, Pirklbauer, Jan, Sach, Marvin, Watanabe, Shinji, Fingscheidt, Tim, Qian, Yanmin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929511784775680
author Zhang, Wangyou
Scheibler, Robin
Saijo, Kohei
Cornell, Samuele
Li, Chenda
Ni, Zhaoheng
Kumar, Anurag
Pirklbauer, Jan
Sach, Marvin
Watanabe, Shinji
Fingscheidt, Tim
Qian, Yanmin
author_facet Zhang, Wangyou
Scheibler, Robin
Saijo, Kohei
Cornell, Samuele
Li, Chenda
Ni, Zhaoheng
Kumar, Anurag
Pirklbauer, Jan
Sach, Marvin
Watanabe, Shinji
Fingscheidt, Tim
Qian, Yanmin
contents The last decade has witnessed significant advancements in deep learning-based speech enhancement (SE). However, most existing SE research has limitations on the coverage of SE sub-tasks, data diversity and amount, and evaluation metrics. To fill this gap and promote research toward universal SE, we establish a new SE challenge, named URGENT, to focus on the universality, robustness, and generalizability of SE. We aim to extend the SE definition to cover different sub-tasks to explore the limits of SE models, starting from denoising, dereverberation, bandwidth extension, and declipping. A novel framework is proposed to unify all these sub-tasks in a single model, allowing the use of all existing SE approaches. We collected public speech and noise data from different domains to construct diverse evaluation data. Finally, we discuss the insights gained from our preliminary baseline experiments based on both generative and discriminative SE methods with 12 curated metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04660
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement
Zhang, Wangyou
Scheibler, Robin
Saijo, Kohei
Cornell, Samuele
Li, Chenda
Ni, Zhaoheng
Kumar, Anurag
Pirklbauer, Jan
Sach, Marvin
Watanabe, Shinji
Fingscheidt, Tim
Qian, Yanmin
Audio and Speech Processing
Sound
The last decade has witnessed significant advancements in deep learning-based speech enhancement (SE). However, most existing SE research has limitations on the coverage of SE sub-tasks, data diversity and amount, and evaluation metrics. To fill this gap and promote research toward universal SE, we establish a new SE challenge, named URGENT, to focus on the universality, robustness, and generalizability of SE. We aim to extend the SE definition to cover different sub-tasks to explore the limits of SE models, starting from denoising, dereverberation, bandwidth extension, and declipping. A novel framework is proposed to unify all these sub-tasks in a single model, allowing the use of all existing SE approaches. We collected public speech and noise data from different domains to construct diverse evaluation data. Finally, we discuss the insights gained from our preliminary baseline experiments based on both generative and discriminative SE methods with 12 curated metrics.
title URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2406.04660