STCON System for the CHiME-8 Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mitrofanov, Anton, Prisyach, Tatiana, Timofeeva, Tatiana, Novoselov, Sergei, Korenevsky, Maxim, Khokhlov, Yuri, Akulov, Artem, Anikin, Alexander, Khalili, Roman, Lezhenin, Iurii, Melnikov, Aleksandr, Miroshnichenko, Dmitriy, Mamaev, Nikita, Odegov, Ilya, Rudnitskaya, Olga, Romanenko, Aleksei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914976389660672
author Mitrofanov, Anton
Prisyach, Tatiana
Timofeeva, Tatiana
Novoselov, Sergei
Korenevsky, Maxim
Khokhlov, Yuri
Akulov, Artem
Anikin, Alexander
Khalili, Roman
Lezhenin, Iurii
Melnikov, Aleksandr
Miroshnichenko, Dmitriy
Mamaev, Nikita
Odegov, Ilya
Rudnitskaya, Olga
Romanenko, Aleksei
author_facet Mitrofanov, Anton
Prisyach, Tatiana
Timofeeva, Tatiana
Novoselov, Sergei
Korenevsky, Maxim
Khokhlov, Yuri
Akulov, Artem
Anikin, Alexander
Khalili, Roman
Lezhenin, Iurii
Melnikov, Aleksandr
Miroshnichenko, Dmitriy
Mamaev, Nikita
Odegov, Ilya
Rudnitskaya, Olga
Romanenko, Aleksei
contents This paper describes the STCON system for the CHiME-8 Challenge Task 1 (DASR) aimed at distant automatic speech transcription and diarization with multiple recording devices. Our main attention was paid to carefully trained and tuned diarization pipeline and speaker counting. This allowed to significantly reduce diarization error rate (DER) and obtain more reliable segments for speech separation and recognition. To improve source separation, we designed a Guided Target speaker Extraction (G-TSE) model and used it in conjunction with the traditional Guided Source Separation (GSS) method. To train various parts of our pipeline, we investigated several data augmentation and generation techniques, which helped us to improve the overall system quality.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13411
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle STCON System for the CHiME-8 Challenge
Mitrofanov, Anton
Prisyach, Tatiana
Timofeeva, Tatiana
Novoselov, Sergei
Korenevsky, Maxim
Khokhlov, Yuri
Akulov, Artem
Anikin, Alexander
Khalili, Roman
Lezhenin, Iurii
Melnikov, Aleksandr
Miroshnichenko, Dmitriy
Mamaev, Nikita
Odegov, Ilya
Rudnitskaya, Olga
Romanenko, Aleksei
Audio and Speech Processing
Sound
This paper describes the STCON system for the CHiME-8 Challenge Task 1 (DASR) aimed at distant automatic speech transcription and diarization with multiple recording devices. Our main attention was paid to carefully trained and tuned diarization pipeline and speaker counting. This allowed to significantly reduce diarization error rate (DER) and obtain more reliable segments for speech separation and recognition. To improve source separation, we designed a Guided Target speaker Extraction (G-TSE) model and used it in conjunction with the traditional Guided Source Separation (GSS) method. To train various parts of our pipeline, we investigated several data augmentation and generation techniques, which helped us to improve the overall system quality.
title STCON System for the CHiME-8 Challenge
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2410.13411