The USTC-NERCSLIP Systems for the CHiME-8 MMCSG Challenge
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909342605770752 |
|---|---|
| author | Jiang, Ya Lan, Hongbo Du, Jun Wang, Qing Niu, Shutong |
| author_facet | Jiang, Ya Lan, Hongbo Du, Jun Wang, Qing Niu, Shutong |
| contents | In the two-person conversation scenario with one wearing smart glasses, transcribing and displaying the speaker's content in real-time is an intriguing application, providing a priori information for subsequent tasks such as translation and comprehension. Meanwhile, multi-modal data captured from the smart glasses is scarce. Therefore, we propose utilizing simulation data with multiple overlap rates and a one-to-one matching training strategy to narrow down the deviation for the model training between real and simulated data. In addition, combining IMU unit data in the model can assist the audio to achieve better real-time speech recognition performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_05986 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | The USTC-NERCSLIP Systems for the CHiME-8 MMCSG Challenge Jiang, Ya Lan, Hongbo Du, Jun Wang, Qing Niu, Shutong Audio and Speech Processing Sound In the two-person conversation scenario with one wearing smart glasses, transcribing and displaying the speaker's content in real-time is an intriguing application, providing a priori information for subsequent tasks such as translation and comprehension. Meanwhile, multi-modal data captured from the smart glasses is scarce. Therefore, we propose utilizing simulation data with multiple overlap rates and a one-to-one matching training strategy to narrow down the deviation for the model training between real and simulated data. In addition, combining IMU unit data in the model can assist the audio to achieve better real-time speech recognition performance. |
| title | The USTC-NERCSLIP Systems for the CHiME-8 MMCSG Challenge |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2410.05986 |