Toward Universal Speech Enhancement for Diverse Input Conditions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Wangyou, Saijo, Kohei, Wang, Zhong-Qiu, Watanabe, Shinji, Qian, Yanmin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913235491356672
author Zhang, Wangyou
Saijo, Kohei
Wang, Zhong-Qiu
Watanabe, Shinji
Qian, Yanmin
author_facet Zhang, Wangyou
Saijo, Kohei
Wang, Zhong-Qiu
Watanabe, Shinji
Qian, Yanmin
contents The past decade has witnessed substantial growth of data-driven speech enhancement (SE) techniques thanks to deep learning. While existing approaches have shown impressive performance in some common datasets, most of them are designed only for a single condition (e.g., single-channel, multi-channel, or a fixed sampling frequency) or only consider a single task (e.g., denoising or dereverberation). Currently, there is no universal SE approach that can effectively handle diverse input conditions with a single model. In this paper, we make the first attempt to investigate this line of research. First, we devise a single SE model that is independent of microphone channels, signal lengths, and sampling frequencies. Second, we design a universal SE benchmark by combining existing public corpora with multiple conditions. Our experiments on a wide range of datasets show that the proposed single model can successfully handle diverse conditions with strong performance.
format Preprint
id arxiv_https___arxiv_org_abs_2309_17384
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Toward Universal Speech Enhancement for Diverse Input Conditions
Zhang, Wangyou
Saijo, Kohei
Wang, Zhong-Qiu
Watanabe, Shinji
Qian, Yanmin
Audio and Speech Processing
Sound
Signal Processing
The past decade has witnessed substantial growth of data-driven speech enhancement (SE) techniques thanks to deep learning. While existing approaches have shown impressive performance in some common datasets, most of them are designed only for a single condition (e.g., single-channel, multi-channel, or a fixed sampling frequency) or only consider a single task (e.g., denoising or dereverberation). Currently, there is no universal SE approach that can effectively handle diverse input conditions with a single model. In this paper, we make the first attempt to investigate this line of research. First, we devise a single SE model that is independent of microphone channels, signal lengths, and sampling frequencies. Second, we design a universal SE benchmark by combining existing public corpora with multiple conditions. Our experiments on a wide range of datasets show that the proposed single model can successfully handle diverse conditions with strong performance.
title Toward Universal Speech Enhancement for Diverse Input Conditions
topic Audio and Speech Processing
Sound
Signal Processing
url https://arxiv.org/abs/2309.17384