Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913242718142464 |
|---|---|
| author | Ge, Zirui Xu, Xinzhou Guo, Haiyan Wang, Tingting Yang, Zhen |
| author_facet | Ge, Zirui Xu, Xinzhou Guo, Haiyan Wang, Tingting Yang, Zhen |
| contents | The emergence of self-supervised representation (i.e., wav2vec 2.0) allows speaker-recognition approaches to process spoken signals through foundation models built on speech data. Nevertheless, effective fusion on the representation requires further investigating, due to the inclusion of fixed or sub-optimal temporal pooling strategies. Despite of improved strategies considering graph learning and graph attention factors, non-injective aggregation still exists in the approaches, which may influence the performance for speaker recognition. In this regard, we propose a speaker recognition approach using Isomorphic Graph ATtention network (IsoGAT) on self-supervised representation. The proposed approach contains three modules of representation learning, graph attention, and aggregation, jointly considering learning on the self-supervised representation and the IsoGAT. Then, we perform experiments for speaker recognition tasks on VoxCeleb1\&2 datasets, with the corresponding experimental results demonstrating the recognition performance for the proposed approach, compared with existing pooling approaches on the self-supervised representation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2308_04666 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation Ge, Zirui Xu, Xinzhou Guo, Haiyan Wang, Tingting Yang, Zhen Sound Audio and Speech Processing The emergence of self-supervised representation (i.e., wav2vec 2.0) allows speaker-recognition approaches to process spoken signals through foundation models built on speech data. Nevertheless, effective fusion on the representation requires further investigating, due to the inclusion of fixed or sub-optimal temporal pooling strategies. Despite of improved strategies considering graph learning and graph attention factors, non-injective aggregation still exists in the approaches, which may influence the performance for speaker recognition. In this regard, we propose a speaker recognition approach using Isomorphic Graph ATtention network (IsoGAT) on self-supervised representation. The proposed approach contains three modules of representation learning, graph attention, and aggregation, jointly considering learning on the self-supervised representation and the IsoGAT. Then, we perform experiments for speaker recognition tasks on VoxCeleb1\&2 datasets, with the corresponding experimental results demonstrating the recognition performance for the proposed approach, compared with existing pooling approaches on the self-supervised representation. |
| title | Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2308.04666 |