MASV: Speaker Verification with Global and Local Context Mamba
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866915063889133568 |
|---|---|
| author | Liu, Yang Wan, Li Huang, Yiteng Sun, Ming Shi, Yangyang Metze, Florian |
| author_facet | Liu, Yang Wan, Li Huang, Yiteng Sun, Ming Shi, Yangyang Metze, Florian |
| contents | Deep learning models like Convolutional Neural Networks and transformers have shown impressive capabilities in speech verification, gaining considerable attention in the research community. However, CNN-based approaches struggle with modeling long-sequence audio effectively, resulting in suboptimal verification performance. On the other hand, transformer-based methods are often hindered by high computational demands, limiting their practicality. This paper presents the MASV model, a novel architecture that integrates the Mamba module into the ECAPA-TDNN framework. By introducing the Local Context Bidirectional Mamba and Tri-Mamba block, the model effectively captures both global and local context within audio sequences. Experimental results demonstrate that the MASV model substantially enhances verification performance, surpassing existing models in both accuracy and efficiency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_10989 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MASV: Speaker Verification with Global and Local Context Mamba Liu, Yang Wan, Li Huang, Yiteng Sun, Ming Shi, Yangyang Metze, Florian Audio and Speech Processing Sound Deep learning models like Convolutional Neural Networks and transformers have shown impressive capabilities in speech verification, gaining considerable attention in the research community. However, CNN-based approaches struggle with modeling long-sequence audio effectively, resulting in suboptimal verification performance. On the other hand, transformer-based methods are often hindered by high computational demands, limiting their practicality. This paper presents the MASV model, a novel architecture that integrates the Mamba module into the ECAPA-TDNN framework. By introducing the Local Context Bidirectional Mamba and Tri-Mamba block, the model effectively captures both global and local context within audio sequences. Experimental results demonstrate that the MASV model substantially enhances verification performance, surpassing existing models in both accuracy and efficiency. |
| title | MASV: Speaker Verification with Global and Local Context Mamba |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2412.10989 |