MASV: Speaker Verification with Global and Local Context Mamba

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Yang, Wan, Li, Huang, Yiteng, Sun, Ming, Shi, Yangyang, Metze, Florian
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915063889133568
author Liu, Yang
Wan, Li
Huang, Yiteng
Sun, Ming
Shi, Yangyang
Metze, Florian
author_facet Liu, Yang
Wan, Li
Huang, Yiteng
Sun, Ming
Shi, Yangyang
Metze, Florian
contents Deep learning models like Convolutional Neural Networks and transformers have shown impressive capabilities in speech verification, gaining considerable attention in the research community. However, CNN-based approaches struggle with modeling long-sequence audio effectively, resulting in suboptimal verification performance. On the other hand, transformer-based methods are often hindered by high computational demands, limiting their practicality. This paper presents the MASV model, a novel architecture that integrates the Mamba module into the ECAPA-TDNN framework. By introducing the Local Context Bidirectional Mamba and Tri-Mamba block, the model effectively captures both global and local context within audio sequences. Experimental results demonstrate that the MASV model substantially enhances verification performance, surpassing existing models in both accuracy and efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10989
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MASV: Speaker Verification with Global and Local Context Mamba
Liu, Yang
Wan, Li
Huang, Yiteng
Sun, Ming
Shi, Yangyang
Metze, Florian
Audio and Speech Processing
Sound
Deep learning models like Convolutional Neural Networks and transformers have shown impressive capabilities in speech verification, gaining considerable attention in the research community. However, CNN-based approaches struggle with modeling long-sequence audio effectively, resulting in suboptimal verification performance. On the other hand, transformer-based methods are often hindered by high computational demands, limiting their practicality. This paper presents the MASV model, a novel architecture that integrates the Mamba module into the ECAPA-TDNN framework. By introducing the Local Context Bidirectional Mamba and Tri-Mamba block, the model effectively captures both global and local context within audio sequences. Experimental results demonstrate that the MASV model substantially enhances verification performance, surpassing existing models in both accuracy and efficiency.
title MASV: Speaker Verification with Global and Local Context Mamba
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2412.10989