A New Perspective on Speaker Verification: Joint Modeling with DFSMN and Transformer

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Hongyu, Li, Hui, Li, Bo
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910593967980544
author Wang, Hongyu
Li, Hui
Li, Bo
author_facet Wang, Hongyu
Li, Hui
Li, Bo
contents Speaker verification is to judge the similarity between two unknown voices in an open set, where the ideal speaker embedding should be able to condense discriminant information into a compact utterance-level representation that has small intra-speaker distances and large inter-speaker distances. We propose Voice Transformer (VOT), a novel model for speaker verification, which integrates parallel transformers at multiple scales. A deep feedforward sequential memory network (DFSMN) is incorporated into the attention part of these transformers to increase feature granularity. The attentive statistics pooling layer is added to focus on important frames and form utterance-level features. We propose Additive Angular Margin Focal Loss (AAMF) to address the hard samples problem. We evaluate the proposed approach on the VoxCeleb1 and CN-Celeb2 datasets, demonstrating that VOT surpasses most mainstream models. The code is available on GitHub\footnote{\url{https://github.com/luckyerr/Voice-Transformer_Speaker-Verification}}.
format Preprint
id arxiv_https___arxiv_org_abs_2312_16826
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A New Perspective on Speaker Verification: Joint Modeling with DFSMN and Transformer
Wang, Hongyu
Li, Hui
Li, Bo
Audio and Speech Processing
Sound
Speaker verification is to judge the similarity between two unknown voices in an open set, where the ideal speaker embedding should be able to condense discriminant information into a compact utterance-level representation that has small intra-speaker distances and large inter-speaker distances. We propose Voice Transformer (VOT), a novel model for speaker verification, which integrates parallel transformers at multiple scales. A deep feedforward sequential memory network (DFSMN) is incorporated into the attention part of these transformers to increase feature granularity. The attentive statistics pooling layer is added to focus on important frames and form utterance-level features. We propose Additive Angular Margin Focal Loss (AAMF) to address the hard samples problem. We evaluate the proposed approach on the VoxCeleb1 and CN-Celeb2 datasets, demonstrating that VOT surpasses most mainstream models. The code is available on GitHub\footnote{\url{https://github.com/luckyerr/Voice-Transformer_Speaker-Verification}}.
title A New Perspective on Speaker Verification: Joint Modeling with DFSMN and Transformer
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2312.16826