A Real-Time Voice Activity Detection Based On Lightweight Neural

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jia, Jidong, Zhao, Pei, Wang, Di
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916261615632384
author Jia, Jidong
Zhao, Pei
Wang, Di
author_facet Jia, Jidong
Zhao, Pei
Wang, Di
contents Voice activity detection (VAD) is the task of detecting speech in an audio stream, which is challenging due to numerous unseen noises and low signal-to-noise ratios in real environments. Recently, neural network-based VADs have alleviated the degradation of performance to some extent. However, the majority of existing studies have employed excessively large models and incorporated future context, while neglecting to evaluate the operational efficiency and latency of the models. In this paper, we propose a lightweight and real-time neural network called MagicNet, which utilizes casual and depth separable 1-D convolutions and GRU. Without relying on future features as input, our proposed model is compared with two state-of-the-art algorithms on synthesized in-domain and out-domain test datasets. The evaluation results demonstrate that MagicNet can achieve improved performance and robustness with fewer parameter costs.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16797
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Real-Time Voice Activity Detection Based On Lightweight Neural
Jia, Jidong
Zhao, Pei
Wang, Di
Sound
Artificial Intelligence
Audio and Speech Processing
Voice activity detection (VAD) is the task of detecting speech in an audio stream, which is challenging due to numerous unseen noises and low signal-to-noise ratios in real environments. Recently, neural network-based VADs have alleviated the degradation of performance to some extent. However, the majority of existing studies have employed excessively large models and incorporated future context, while neglecting to evaluate the operational efficiency and latency of the models. In this paper, we propose a lightweight and real-time neural network called MagicNet, which utilizes casual and depth separable 1-D convolutions and GRU. Without relying on future features as input, our proposed model is compared with two state-of-the-art algorithms on synthesized in-domain and out-domain test datasets. The evaluation results demonstrate that MagicNet can achieve improved performance and robustness with fewer parameter costs.
title A Real-Time Voice Activity Detection Based On Lightweight Neural
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2405.16797