GMM-ResNext: Combining Generative and Discriminative Models for Speaker Verification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Hui, Lei, Zhenchun, Liu, Changhong, Zhou, Yong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909241012387840
author Yan, Hui
Lei, Zhenchun
Liu, Changhong
Zhou, Yong
author_facet Yan, Hui
Lei, Zhenchun
Liu, Changhong
Zhou, Yong
contents With the development of deep learning, many different network architectures have been explored in speaker verification. However, most network architectures rely on a single deep learning architecture, and hybrid networks combining different architectures have been little studied in ASV tasks. In this paper, we propose the GMM-ResNext model for speaker verification. Conventional GMM does not consider the score distribution of each frame feature over all Gaussian components and ignores the relationship between neighboring speech frames. So, we extract the log Gaussian probability features based on the raw acoustic features and use ResNext-based network as the backbone to extract the speaker embedding. GMM-ResNext combines Generative and Discriminative Models to improve the generalization ability of deep learning models and allows one to more easily specify meaningful priors on model parameters. A two-path GMM-ResNext model based on two gender-related GMMs has also been proposed. The Experimental results show that the proposed GMM-ResNext achieves relative improvements of 48.1\% and 11.3\% in EER compared with ResNet34 and ECAPA-TDNN on VoxCeleb1-O test set.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03135
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GMM-ResNext: Combining Generative and Discriminative Models for Speaker Verification
Yan, Hui
Lei, Zhenchun
Liu, Changhong
Zhou, Yong
Sound
Artificial Intelligence
Human-Computer Interaction
Audio and Speech Processing
With the development of deep learning, many different network architectures have been explored in speaker verification. However, most network architectures rely on a single deep learning architecture, and hybrid networks combining different architectures have been little studied in ASV tasks. In this paper, we propose the GMM-ResNext model for speaker verification. Conventional GMM does not consider the score distribution of each frame feature over all Gaussian components and ignores the relationship between neighboring speech frames. So, we extract the log Gaussian probability features based on the raw acoustic features and use ResNext-based network as the backbone to extract the speaker embedding. GMM-ResNext combines Generative and Discriminative Models to improve the generalization ability of deep learning models and allows one to more easily specify meaningful priors on model parameters. A two-path GMM-ResNext model based on two gender-related GMMs has also been proposed. The Experimental results show that the proposed GMM-ResNext achieves relative improvements of 48.1\% and 11.3\% in EER compared with ResNet34 and ECAPA-TDNN on VoxCeleb1-O test set.
title GMM-ResNext: Combining Generative and Discriminative Models for Speaker Verification
topic Sound
Artificial Intelligence
Human-Computer Interaction
Audio and Speech Processing
url https://arxiv.org/abs/2407.03135