LightCAM: A Fast and Light Implementation of Context-Aware Masking based D-TDNN for Speaker Verification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cao, Di, Wang, Xianchen, Zhou, Junfeng, Zhang, Jiakai, Lei, Yanjing, Chen, Wenpeng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916122483228672
author Cao, Di
Wang, Xianchen
Zhou, Junfeng
Zhang, Jiakai
Lei, Yanjing
Chen, Wenpeng
author_facet Cao, Di
Wang, Xianchen
Zhou, Junfeng
Zhang, Jiakai
Lei, Yanjing
Chen, Wenpeng
contents Traditional Time Delay Neural Networks (TDNN) have achieved state-of-the-art performance at the cost of high computational complexity and slower inference speed, making them difficult to implement in an industrial environment. The Densely Connected Time Delay Neural Network (D-TDNN) with Context Aware Masking (CAM) module has proven to be an efficient structure to reduce complexity while maintaining system performance. In this paper, we propose a fast and lightweight model, LightCAM, which further adopts a depthwise separable convolution module (DSM) and uses multi-scale feature aggregation (MFA) for feature fusion at different levels. Extensive experiments are conducted on VoxCeleb dataset, the comparative results show that it has achieved an EER of 0.83 and MinDCF of 0.0891 in VoxCeleb1-O, which outperforms the other mainstream speaker verification methods. In addition, complexity analysis further demonstrates that the proposed architecture has lower computational cost and faster inference speed.
format Preprint
id arxiv_https___arxiv_org_abs_2402_06073
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LightCAM: A Fast and Light Implementation of Context-Aware Masking based D-TDNN for Speaker Verification
Cao, Di
Wang, Xianchen
Zhou, Junfeng
Zhang, Jiakai
Lei, Yanjing
Chen, Wenpeng
Computation and Language
Sound
Audio and Speech Processing
Traditional Time Delay Neural Networks (TDNN) have achieved state-of-the-art performance at the cost of high computational complexity and slower inference speed, making them difficult to implement in an industrial environment. The Densely Connected Time Delay Neural Network (D-TDNN) with Context Aware Masking (CAM) module has proven to be an efficient structure to reduce complexity while maintaining system performance. In this paper, we propose a fast and lightweight model, LightCAM, which further adopts a depthwise separable convolution module (DSM) and uses multi-scale feature aggregation (MFA) for feature fusion at different levels. Extensive experiments are conducted on VoxCeleb dataset, the comparative results show that it has achieved an EER of 0.83 and MinDCF of 0.0891 in VoxCeleb1-O, which outperforms the other mainstream speaker verification methods. In addition, complexity analysis further demonstrates that the proposed architecture has lower computational cost and faster inference speed.
title LightCAM: A Fast and Light Implementation of Context-Aware Masking based D-TDNN for Speaker Verification
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2402.06073