TCG CREST System Description for the Second DISPLACE Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Raghav, Nikhil, Saha, Subhajit, Sahidullah, Md, Das, Swagatam
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910617887047680
author Raghav, Nikhil
Saha, Subhajit
Sahidullah, Md
Das, Swagatam
author_facet Raghav, Nikhil
Saha, Subhajit
Sahidullah, Md
Das, Swagatam
contents In this report, we describe the speaker diarization (SD) and language diarization (LD) systems developed by our team for the Second DISPLACE Challenge, 2024. Our contributions were dedicated to Track 1 for SD and Track 2 for LD in multilingual and multi-speaker scenarios. We investigated different speech enhancement techniques, voice activity detection (VAD) techniques, unsupervised domain categorization, and neural embedding extraction architectures. We also exploited the fusion of various embedding extraction models. We implemented our system with the open-source SpeechBrain toolkit. Our final submissions use spectral clustering for both the speaker and language diarization. We achieve about $7\%$ relative improvement over the challenge baseline in Track 1. We did not obtain improvement over the challenge baseline in Track 2.
format Preprint
id arxiv_https___arxiv_org_abs_2409_15356
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TCG CREST System Description for the Second DISPLACE Challenge
Raghav, Nikhil
Saha, Subhajit
Sahidullah, Md
Das, Swagatam
Audio and Speech Processing
Machine Learning
Sound
In this report, we describe the speaker diarization (SD) and language diarization (LD) systems developed by our team for the Second DISPLACE Challenge, 2024. Our contributions were dedicated to Track 1 for SD and Track 2 for LD in multilingual and multi-speaker scenarios. We investigated different speech enhancement techniques, voice activity detection (VAD) techniques, unsupervised domain categorization, and neural embedding extraction architectures. We also exploited the fusion of various embedding extraction models. We implemented our system with the open-source SpeechBrain toolkit. Our final submissions use spectral clustering for both the speaker and language diarization. We achieve about $7\%$ relative improvement over the challenge baseline in Track 1. We did not obtain improvement over the challenge baseline in Track 2.
title TCG CREST System Description for the Second DISPLACE Challenge
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2409.15356