Robust Channel Learning for Large-Scale Radio Speaker Verification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Wenhao, Wei, Jianguo, Lu, Wenhuan, Li, Lei, Lu, Xugang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916288904822784
author Yang, Wenhao
Wei, Jianguo
Lu, Wenhuan
Li, Lei
Lu, Xugang
author_facet Yang, Wenhao
Wei, Jianguo
Lu, Wenhuan
Li, Lei
Lu, Xugang
contents Recent research in speaker verification has increasingly focused on achieving robust and reliable recognition under challenging channel conditions and noisy environments. Identifying speakers in radio communications is particularly difficult due to inherent limitations such as constrained bandwidth and pervasive noise interference. To address this issue, we present a Channel Robust Speaker Learning (CRSL) framework that enhances the robustness of the current speaker verification pipeline, considering data source, data augmentation, and the efficiency of model transfer processes. Our framework introduces an augmentation module that mitigates bandwidth variations in radio speech datasets by manipulating the bandwidth of training inputs. It also addresses unknown noise by introducing noise within the manifold space. Additionally, we propose an efficient fine-tuning method that reduces the need for extensive additional training time and large amounts of data. Moreover, we develop a toolkit for assembling a large-scale radio speech corpus and establish a benchmark specifically tailored for radio scenario speaker verification studies. Experimental results demonstrate that our proposed methodology effectively enhances performance and mitigates degradation caused by radio transmission in speaker verification tasks. The code will be available on Github.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10956
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Robust Channel Learning for Large-Scale Radio Speaker Verification
Yang, Wenhao
Wei, Jianguo
Lu, Wenhuan
Li, Lei
Lu, Xugang
Sound
Machine Learning
Audio and Speech Processing
Recent research in speaker verification has increasingly focused on achieving robust and reliable recognition under challenging channel conditions and noisy environments. Identifying speakers in radio communications is particularly difficult due to inherent limitations such as constrained bandwidth and pervasive noise interference. To address this issue, we present a Channel Robust Speaker Learning (CRSL) framework that enhances the robustness of the current speaker verification pipeline, considering data source, data augmentation, and the efficiency of model transfer processes. Our framework introduces an augmentation module that mitigates bandwidth variations in radio speech datasets by manipulating the bandwidth of training inputs. It also addresses unknown noise by introducing noise within the manifold space. Additionally, we propose an efficient fine-tuning method that reduces the need for extensive additional training time and large amounts of data. Moreover, we develop a toolkit for assembling a large-scale radio speech corpus and establish a benchmark specifically tailored for radio scenario speaker verification studies. Experimental results demonstrate that our proposed methodology effectively enhances performance and mitigates degradation caused by radio transmission in speaker verification tasks. The code will be available on Github.
title Robust Channel Learning for Large-Scale Radio Speaker Verification
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2406.10956