STCON System for the CHiME-8 Challenge
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914976389660672 |
|---|---|
| author | Mitrofanov, Anton Prisyach, Tatiana Timofeeva, Tatiana Novoselov, Sergei Korenevsky, Maxim Khokhlov, Yuri Akulov, Artem Anikin, Alexander Khalili, Roman Lezhenin, Iurii Melnikov, Aleksandr Miroshnichenko, Dmitriy Mamaev, Nikita Odegov, Ilya Rudnitskaya, Olga Romanenko, Aleksei |
| author_facet | Mitrofanov, Anton Prisyach, Tatiana Timofeeva, Tatiana Novoselov, Sergei Korenevsky, Maxim Khokhlov, Yuri Akulov, Artem Anikin, Alexander Khalili, Roman Lezhenin, Iurii Melnikov, Aleksandr Miroshnichenko, Dmitriy Mamaev, Nikita Odegov, Ilya Rudnitskaya, Olga Romanenko, Aleksei |
| contents | This paper describes the STCON system for the CHiME-8 Challenge Task 1 (DASR) aimed at distant automatic speech transcription and diarization with multiple recording devices. Our main attention was paid to carefully trained and tuned diarization pipeline and speaker counting. This allowed to significantly reduce diarization error rate (DER) and obtain more reliable segments for speech separation and recognition. To improve source separation, we designed a Guided Target speaker Extraction (G-TSE) model and used it in conjunction with the traditional Guided Source Separation (GSS) method. To train various parts of our pipeline, we investigated several data augmentation and generation techniques, which helped us to improve the overall system quality. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_13411 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | STCON System for the CHiME-8 Challenge Mitrofanov, Anton Prisyach, Tatiana Timofeeva, Tatiana Novoselov, Sergei Korenevsky, Maxim Khokhlov, Yuri Akulov, Artem Anikin, Alexander Khalili, Roman Lezhenin, Iurii Melnikov, Aleksandr Miroshnichenko, Dmitriy Mamaev, Nikita Odegov, Ilya Rudnitskaya, Olga Romanenko, Aleksei Audio and Speech Processing Sound This paper describes the STCON system for the CHiME-8 Challenge Task 1 (DASR) aimed at distant automatic speech transcription and diarization with multiple recording devices. Our main attention was paid to carefully trained and tuned diarization pipeline and speaker counting. This allowed to significantly reduce diarization error rate (DER) and obtain more reliable segments for speech separation and recognition. To improve source separation, we designed a Guided Target speaker Extraction (G-TSE) model and used it in conjunction with the traditional Guided Source Separation (GSS) method. To train various parts of our pipeline, we investigated several data augmentation and generation techniques, which helped us to improve the overall system quality. |
| title | STCON System for the CHiME-8 Challenge |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2410.13411 |