Team HYU ASML ROBOVOX SP Cup 2024 System Description

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Jeong-Hwan, Kim, Gaeun, Lee, Hee-Jae, Ahn, Seyun, Kim, Hyun-Soo, Chang, Joon-Hyuk
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911956790673408
author Choi, Jeong-Hwan
Kim, Gaeun
Lee, Hee-Jae
Ahn, Seyun
Kim, Hyun-Soo
Chang, Joon-Hyuk
author_facet Choi, Jeong-Hwan
Kim, Gaeun
Lee, Hee-Jae
Ahn, Seyun
Kim, Hyun-Soo
Chang, Joon-Hyuk
contents This report describes the submission of HYU ASML team to the IEEE Signal Processing Cup 2024 (SP Cup 2024). This challenge, titled "ROBOVOX: Far-Field Speaker Recognition by a Mobile Robot," focuses on speaker recognition using a mobile robot in noisy and reverberant conditions. Our solution combines the result of deep residual neural networks and time-delay neural network-based speaker embedding models. These models were trained on a diverse dataset that includes French speech. To account for the challenging evaluation environment characterized by high noise, reverberation, and short speech conditions, we focused on data augmentation and training speech duration for the speaker embedding model. Our submission achieved second place on the SP Cup 2024 public leaderboard, with a detection cost function of 0.5245 and an equal error rate of 6.46%.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11365
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Team HYU ASML ROBOVOX SP Cup 2024 System Description
Choi, Jeong-Hwan
Kim, Gaeun
Lee, Hee-Jae
Ahn, Seyun
Kim, Hyun-Soo
Chang, Joon-Hyuk
Audio and Speech Processing
This report describes the submission of HYU ASML team to the IEEE Signal Processing Cup 2024 (SP Cup 2024). This challenge, titled "ROBOVOX: Far-Field Speaker Recognition by a Mobile Robot," focuses on speaker recognition using a mobile robot in noisy and reverberant conditions. Our solution combines the result of deep residual neural networks and time-delay neural network-based speaker embedding models. These models were trained on a diverse dataset that includes French speech. To account for the challenging evaluation environment characterized by high noise, reverberation, and short speech conditions, we focused on data augmentation and training speech duration for the speaker embedding model. Our submission achieved second place on the SP Cup 2024 public leaderboard, with a detection cost function of 0.5245 and an equal error rate of 6.46%.
title Team HYU ASML ROBOVOX SP Cup 2024 System Description
topic Audio and Speech Processing
url https://arxiv.org/abs/2407.11365