TSE-PI: Target Sound Extraction under Reverberant Environments with Pitch Information

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yiwen, Wu, Xihong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913389953941504
author Wang, Yiwen
Wu, Xihong
author_facet Wang, Yiwen
Wu, Xihong
contents Target sound extraction (TSE) separates the target sound from the mixture signals based on provided clues. However, the performance of existing models significantly degrades under reverberant conditions. Inspired by auditory scene analysis (ASA), this work proposes a TSE model provided with pitch information named TSE-PI. Conditional pitch extraction is achieved through the Feature-wise Linearly Modulated layer with the sound-class label. A modified Waveformer model combined with pitch information, employing a learnable Gammatone filterbank in place of the convolutional encoder, is used for target sound extraction. The inclusion of pitch information is aimed at improving the model's performance. The experimental results on the FSD50K dataset illustrate 2.4 dB improvements of target sound extraction under reverberant environments when incorporating pitch information and Gammatone filterbank.
format Preprint
id arxiv_https___arxiv_org_abs_2406_08716
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TSE-PI: Target Sound Extraction under Reverberant Environments with Pitch Information
Wang, Yiwen
Wu, Xihong
Sound
Audio and Speech Processing
Target sound extraction (TSE) separates the target sound from the mixture signals based on provided clues. However, the performance of existing models significantly degrades under reverberant conditions. Inspired by auditory scene analysis (ASA), this work proposes a TSE model provided with pitch information named TSE-PI. Conditional pitch extraction is achieved through the Feature-wise Linearly Modulated layer with the sound-class label. A modified Waveformer model combined with pitch information, employing a learnable Gammatone filterbank in place of the convolutional encoder, is used for target sound extraction. The inclusion of pitch information is aimed at improving the model's performance. The experimental results on the FSD50K dataset illustrate 2.4 dB improvements of target sound extraction under reverberant environments when incorporating pitch information and Gammatone filterbank.
title TSE-PI: Target Sound Extraction under Reverberant Environments with Pitch Information
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.08716