LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Menon, Aditya Srinivas, Gohil, Raj Prakash, Tripathi, Kumud, Wasnik, Pankaj
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913871975940096
author Menon, Aditya Srinivas
Gohil, Raj Prakash
Tripathi, Kumud
Wasnik, Pankaj
author_facet Menon, Aditya Srinivas
Gohil, Raj Prakash
Tripathi, Kumud
Wasnik, Pankaj
contents Speaker recognition models face challenges in multi-lingual settings due to the entanglement of linguistic information within speaker embeddings. The overlap between vocal traits such as accent, vocal anatomy, and a language's phonetic structure complicates separating linguistic and speaker information. Disentangling these components can significantly improve speaker recognition accuracy. To this end, we propose a novel disentanglement learning strategy that integrates joint learning through prefix-tuned cross-attention. This approach is particularly effective when speakers switch between languages. Experimental results show the model generalizes across monolingual and multi-lingual settings, including unseen languages. Notably, the proposed model improves the equal error rate across multiple datasets, highlighting its ability to separate language information from speaker embeddings and enhance recognition in diverse linguistic conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02083
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
Menon, Aditya Srinivas
Gohil, Raj Prakash
Tripathi, Kumud
Wasnik, Pankaj
Sound
Artificial Intelligence
Machine Learning
Multimedia
Speaker recognition models face challenges in multi-lingual settings due to the entanglement of linguistic information within speaker embeddings. The overlap between vocal traits such as accent, vocal anatomy, and a language's phonetic structure complicates separating linguistic and speaker information. Disentangling these components can significantly improve speaker recognition accuracy. To this end, we propose a novel disentanglement learning strategy that integrates joint learning through prefix-tuned cross-attention. This approach is particularly effective when speakers switch between languages. Experimental results show the model generalizes across monolingual and multi-lingual settings, including unseen languages. Notably, the proposed model improves the equal error rate across multiple datasets, highlighting its ability to separate language information from speaker embeddings and enhance recognition in diverse linguistic conditions.
title LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
topic Sound
Artificial Intelligence
Machine Learning
Multimedia
url https://arxiv.org/abs/2506.02083