ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jung, Jee-weon, Zhang, Wangyou, Shi, Jiatong, Aldeneh, Zakaria, Higuchi, Takuya, Theobald, Barry-John, Abdelaziz, Ahmed Hussen, Watanabe, Shinji
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914833275813888
author Jung, Jee-weon
Zhang, Wangyou
Shi, Jiatong
Aldeneh, Zakaria
Higuchi, Takuya
Theobald, Barry-John
Abdelaziz, Ahmed Hussen
Watanabe, Shinji
author_facet Jung, Jee-weon
Zhang, Wangyou
Shi, Jiatong
Aldeneh, Zakaria
Higuchi, Takuya
Theobald, Barry-John
Abdelaziz, Ahmed Hussen
Watanabe, Shinji
contents This paper introduces ESPnet-SPK, a toolkit designed with several objectives for training speaker embedding extractors. First, we provide an open-source platform for researchers in the speaker recognition community to effortlessly build models. We provide several models, ranging from x-vector to recent SKA-TDNN. Through the modularized architecture design, variants can be developed easily. We also aspire to bridge developed models with other domains, facilitating the broad research community to effortlessly incorporate state-of-the-art embedding extractors. Pre-trained embedding extractors can be accessed in an off-the-shelf manner and we demonstrate the toolkit's versatility by showcasing its integration with two tasks. Another goal is to integrate with diverse self-supervised learning features. We release a reproducible recipe that achieves an equal error rate of 0.39% on the Vox1-O evaluation protocol using WavLM-Large with ECAPA-TDNN.
format Preprint
id arxiv_https___arxiv_org_abs_2401_17230
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
Jung, Jee-weon
Zhang, Wangyou
Shi, Jiatong
Aldeneh, Zakaria
Higuchi, Takuya
Theobald, Barry-John
Abdelaziz, Ahmed Hussen
Watanabe, Shinji
Sound
Artificial Intelligence
Audio and Speech Processing
This paper introduces ESPnet-SPK, a toolkit designed with several objectives for training speaker embedding extractors. First, we provide an open-source platform for researchers in the speaker recognition community to effortlessly build models. We provide several models, ranging from x-vector to recent SKA-TDNN. Through the modularized architecture design, variants can be developed easily. We also aspire to bridge developed models with other domains, facilitating the broad research community to effortlessly incorporate state-of-the-art embedding extractors. Pre-trained embedding extractors can be accessed in an off-the-shelf manner and we demonstrate the toolkit's versatility by showcasing its integration with two tasks. Another goal is to integrate with diverse self-supervised learning features. We release a reproducible recipe that achieves an equal error rate of 0.39% on the Vox1-O evaluation protocol using WavLM-Large with ECAPA-TDNN.
title ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2401.17230