Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhao, Yiyang, Wang, Shuai, Sun, Guangzhi, Chen, Zehua, Zhang, Chao, Xu, Mingxing, Zheng, Thomas Fang
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917761357185024
author Zhao, Yiyang
Wang, Shuai
Sun, Guangzhi
Chen, Zehua
Zhang, Chao
Xu, Mingxing
Zheng, Thomas Fang
author_facet Zhao, Yiyang
Wang, Shuai
Sun, Guangzhi
Chen, Zehua
Zhang, Chao
Xu, Mingxing
Zheng, Thomas Fang
contents In this paper, Whisper, a large-scale pre-trained model for automatic speech recognition, is proposed to apply to speaker verification. A partial multi-scale feature aggregation (PMFA) approach is proposed based on a subset of Whisper encoder blocks to derive highly discriminative speaker embeddings.Experimental results demonstrate that using the middle to later blocks of the Whisper encoder keeps more speaker information. On the VoxCeleb1 and CN-Celeb1 datasets, our system achieves 1.42% and 8.23% equal error rates (EERs) respectively, receiving 0.58% and 1.81% absolute EER reductions over the ECAPA-TDNN baseline, and 0.46% and 0.97% over the ResNet34 baseline. Furthermore, our results indicate that using Whisper models trained on multilingual data can effectively enhance the model's robustness across languages. Finally, the low-rank adaptation approach is evaluated, which reduces the trainable model parameters by approximately 45 times while only slightly increasing EER by 0.2%.
format Preprint
id arxiv_https___arxiv_org_abs_2408_15585
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
Zhao, Yiyang
Wang, Shuai
Sun, Guangzhi
Chen, Zehua
Zhang, Chao
Xu, Mingxing
Zheng, Thomas Fang
Sound
Audio and Speech Processing
In this paper, Whisper, a large-scale pre-trained model for automatic speech recognition, is proposed to apply to speaker verification. A partial multi-scale feature aggregation (PMFA) approach is proposed based on a subset of Whisper encoder blocks to derive highly discriminative speaker embeddings.Experimental results demonstrate that using the middle to later blocks of the Whisper encoder keeps more speaker information. On the VoxCeleb1 and CN-Celeb1 datasets, our system achieves 1.42% and 8.23% equal error rates (EERs) respectively, receiving 0.58% and 1.81% absolute EER reductions over the ECAPA-TDNN baseline, and 0.46% and 0.97% over the ResNet34 baseline. Furthermore, our results indicate that using Whisper models trained on multilingual data can effectively enhance the model's robustness across languages. Finally, the low-rank adaptation approach is evaluated, which reduces the trainable model parameters by approximately 45 times while only slightly increasing EER by 0.2%.
title Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2408.15585