Cross-Referencing Self-Training Network for Sound Event Detection in Audio Mixtures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Sangwook, Han, David K., Elhilali, Mounya
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909442421817344
author Park, Sangwook
Han, David K.
Elhilali, Mounya
author_facet Park, Sangwook
Han, David K.
Elhilali, Mounya
contents Sound event detection is an important facet of audio tagging that aims to identify sounds of interest and define both the sound category and time boundaries for each sound event in a continuous recording. With advances in deep neural networks, there has been tremendous improvement in the performance of sound event detection systems, although at the expense of costly data collection and labeling efforts. In fact, current state-of-the-art methods employ supervised training methods that leverage large amounts of data samples and corresponding labels in order to facilitate identification of sound category and time stamps of events. As an alternative, the current study proposes a semi-supervised method for generating pseudo-labels from unsupervised data using a student-teacher scheme that balances self-training and cross-training. Additionally, this paper explores post-processing which extracts sound intervals from network prediction, for further improvement in sound event detection performance. The proposed approach is evaluated on sound event detection task for the DCASE2020 challenge. The results of these methods on both "validation" and "public evaluation" sets of DESED database show significant improvement compared to the state-of-the art systems in semi-supervised learning.
format Preprint
id arxiv_https___arxiv_org_abs_2105_13392
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Cross-Referencing Self-Training Network for Sound Event Detection in Audio Mixtures
Park, Sangwook
Han, David K.
Elhilali, Mounya
Sound
Machine Learning
Audio and Speech Processing
Sound event detection is an important facet of audio tagging that aims to identify sounds of interest and define both the sound category and time boundaries for each sound event in a continuous recording. With advances in deep neural networks, there has been tremendous improvement in the performance of sound event detection systems, although at the expense of costly data collection and labeling efforts. In fact, current state-of-the-art methods employ supervised training methods that leverage large amounts of data samples and corresponding labels in order to facilitate identification of sound category and time stamps of events. As an alternative, the current study proposes a semi-supervised method for generating pseudo-labels from unsupervised data using a student-teacher scheme that balances self-training and cross-training. Additionally, this paper explores post-processing which extracts sound intervals from network prediction, for further improvement in sound event detection performance. The proposed approach is evaluated on sound event detection task for the DCASE2020 challenge. The results of these methods on both "validation" and "public evaluation" sets of DESED database show significant improvement compared to the state-of-the art systems in semi-supervised learning.
title Cross-Referencing Self-Training Network for Sound Event Detection in Audio Mixtures
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2105.13392