TCG CREST System Description for the Second DISPLACE Challenge
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910617887047680 |
|---|---|
| author | Raghav, Nikhil Saha, Subhajit Sahidullah, Md Das, Swagatam |
| author_facet | Raghav, Nikhil Saha, Subhajit Sahidullah, Md Das, Swagatam |
| contents | In this report, we describe the speaker diarization (SD) and language diarization (LD) systems developed by our team for the Second DISPLACE Challenge, 2024. Our contributions were dedicated to Track 1 for SD and Track 2 for LD in multilingual and multi-speaker scenarios. We investigated different speech enhancement techniques, voice activity detection (VAD) techniques, unsupervised domain categorization, and neural embedding extraction architectures. We also exploited the fusion of various embedding extraction models. We implemented our system with the open-source SpeechBrain toolkit. Our final submissions use spectral clustering for both the speaker and language diarization. We achieve about $7\%$ relative improvement over the challenge baseline in Track 1. We did not obtain improvement over the challenge baseline in Track 2. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_15356 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | TCG CREST System Description for the Second DISPLACE Challenge Raghav, Nikhil Saha, Subhajit Sahidullah, Md Das, Swagatam Audio and Speech Processing Machine Learning Sound In this report, we describe the speaker diarization (SD) and language diarization (LD) systems developed by our team for the Second DISPLACE Challenge, 2024. Our contributions were dedicated to Track 1 for SD and Track 2 for LD in multilingual and multi-speaker scenarios. We investigated different speech enhancement techniques, voice activity detection (VAD) techniques, unsupervised domain categorization, and neural embedding extraction architectures. We also exploited the fusion of various embedding extraction models. We implemented our system with the open-source SpeechBrain toolkit. Our final submissions use spectral clustering for both the speaker and language diarization. We achieve about $7\%$ relative improvement over the challenge baseline in Track 1. We did not obtain improvement over the challenge baseline in Track 2. |
| title | TCG CREST System Description for the Second DISPLACE Challenge |
| topic | Audio and Speech Processing Machine Learning Sound |
| url | https://arxiv.org/abs/2409.15356 |