ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thienpondt, Jenthe, Demuynck, Kris
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909074914803712
author Thienpondt, Jenthe
Demuynck, Kris
author_facet Thienpondt, Jenthe
Demuynck, Kris
contents In this paper, we present ECAPA2, a novel hybrid neural network architecture and training strategy to produce robust speaker embeddings. Most speaker verification models are based on either the 1D- or 2D-convolutional operation, often manifested as Time Delay Neural Networks or ResNets, respectively. Hybrid models are relatively unexplored without an intuitive explanation what constitutes best practices in regard to its architectural choices. We motivate the proposed ECAPA2 model in this paper with an analysis of current speaker verification architectures. In addition, we propose a training strategy which makes the speaker embeddings more robust against overlapping speech and short utterance lengths. The presented ECAPA2 architecture and training strategy attains state-of-the-art performance on the VoxCeleb1 test sets with significantly less parameters than current models. Finally, we make a pre-trained model publicly available to promote research on downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2401_08342
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker Embeddings
Thienpondt, Jenthe
Demuynck, Kris
Audio and Speech Processing
In this paper, we present ECAPA2, a novel hybrid neural network architecture and training strategy to produce robust speaker embeddings. Most speaker verification models are based on either the 1D- or 2D-convolutional operation, often manifested as Time Delay Neural Networks or ResNets, respectively. Hybrid models are relatively unexplored without an intuitive explanation what constitutes best practices in regard to its architectural choices. We motivate the proposed ECAPA2 model in this paper with an analysis of current speaker verification architectures. In addition, we propose a training strategy which makes the speaker embeddings more robust against overlapping speech and short utterance lengths. The presented ECAPA2 architecture and training strategy attains state-of-the-art performance on the VoxCeleb1 test sets with significantly less parameters than current models. Finally, we make a pre-trained model publicly available to promote research on downstream tasks.
title ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker Embeddings
topic Audio and Speech Processing
url https://arxiv.org/abs/2401.08342