Saved in:
Bibliographic Details
Main Authors: Bovbjerg, Holger Severin, Jensen, Jesper, Østergaard, Jan, Tan, Zheng-Hua
Format: Preprint
Published: 2023
Subjects:
Online Access:https://arxiv.org/abs/2312.16613
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913204906491904
author Bovbjerg, Holger Severin
Jensen, Jesper
Østergaard, Jan
Tan, Zheng-Hua
author_facet Bovbjerg, Holger Severin
Jensen, Jesper
Østergaard, Jan
Tan, Zheng-Hua
contents In this paper, we propose the use of self-supervised pretraining on a large unlabelled data set to improve the performance of a personalized voice activity detection (VAD) model in adverse conditions. We pretrain a long short-term memory (LSTM)-encoder using the autoregressive predictive coding (APC) framework and fine-tune it for personalized VAD. We also propose a denoising variant of APC, with the goal of improving the robustness of personalized VAD. The trained models are systematically evaluated on both clean speech and speech contaminated by various types of noise at different SNR-levels and compared to a purely supervised model. Our experiments show that self-supervised pretraining not only improves performance in clean conditions, but also yields models which are more robust to adverse conditions compared to purely supervised learning.
format Preprint
id arxiv_https___arxiv_org_abs_2312_16613
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Self-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
Bovbjerg, Holger Severin
Jensen, Jesper
Østergaard, Jan
Tan, Zheng-Hua
Sound
Machine Learning
Audio and Speech Processing
68T10
I.2.6
In this paper, we propose the use of self-supervised pretraining on a large unlabelled data set to improve the performance of a personalized voice activity detection (VAD) model in adverse conditions. We pretrain a long short-term memory (LSTM)-encoder using the autoregressive predictive coding (APC) framework and fine-tune it for personalized VAD. We also propose a denoising variant of APC, with the goal of improving the robustness of personalized VAD. The trained models are systematically evaluated on both clean speech and speech contaminated by various types of noise at different SNR-levels and compared to a purely supervised model. Our experiments show that self-supervised pretraining not only improves performance in clean conditions, but also yields models which are more robust to adverse conditions compared to purely supervised learning.
title Self-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
topic Sound
Machine Learning
Audio and Speech Processing
68T10
I.2.6
url https://arxiv.org/abs/2312.16613