Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ge, Zirui, Xu, Xinzhou, Guo, Haiyan, Wang, Tingting, Yang, Zhen
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913242718142464
author Ge, Zirui
Xu, Xinzhou
Guo, Haiyan
Wang, Tingting
Yang, Zhen
author_facet Ge, Zirui
Xu, Xinzhou
Guo, Haiyan
Wang, Tingting
Yang, Zhen
contents The emergence of self-supervised representation (i.e., wav2vec 2.0) allows speaker-recognition approaches to process spoken signals through foundation models built on speech data. Nevertheless, effective fusion on the representation requires further investigating, due to the inclusion of fixed or sub-optimal temporal pooling strategies. Despite of improved strategies considering graph learning and graph attention factors, non-injective aggregation still exists in the approaches, which may influence the performance for speaker recognition. In this regard, we propose a speaker recognition approach using Isomorphic Graph ATtention network (IsoGAT) on self-supervised representation. The proposed approach contains three modules of representation learning, graph attention, and aggregation, jointly considering learning on the self-supervised representation and the IsoGAT. Then, we perform experiments for speaker recognition tasks on VoxCeleb1\&2 datasets, with the corresponding experimental results demonstrating the recognition performance for the proposed approach, compared with existing pooling approaches on the self-supervised representation.
format Preprint
id arxiv_https___arxiv_org_abs_2308_04666
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation
Ge, Zirui
Xu, Xinzhou
Guo, Haiyan
Wang, Tingting
Yang, Zhen
Sound
Audio and Speech Processing
The emergence of self-supervised representation (i.e., wav2vec 2.0) allows speaker-recognition approaches to process spoken signals through foundation models built on speech data. Nevertheless, effective fusion on the representation requires further investigating, due to the inclusion of fixed or sub-optimal temporal pooling strategies. Despite of improved strategies considering graph learning and graph attention factors, non-injective aggregation still exists in the approaches, which may influence the performance for speaker recognition. In this regard, we propose a speaker recognition approach using Isomorphic Graph ATtention network (IsoGAT) on self-supervised representation. The proposed approach contains three modules of representation learning, graph attention, and aggregation, jointly considering learning on the self-supervised representation and the IsoGAT. Then, we perform experiments for speaker recognition tasks on VoxCeleb1\&2 datasets, with the corresponding experimental results demonstrating the recognition performance for the proposed approach, compared with existing pooling approaches on the self-supervised representation.
title Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2308.04666