A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909414849511424 |
|---|---|
| author | Song, Zheshu Ma, Ziyang Yang, Yifan Zhuo, Jianheng Chen, Xie |
| author_facet | Song, Zheshu Ma, Ziyang Yang, Yifan Zhuo, Jianheng Chen, Xie |
| contents | Large Language Models (LLMs) have showcased exceptional performance across diverse NLP tasks, and their integration with speech encoder is rapidly emerging as a dominant trend in the Automatic Speech Recognition (ASR) field. Previous works mainly concentrated on leveraging LLMs for speech recognition in English and Chinese. However, their potential for addressing speech recognition challenges in low resource settings remains underexplored. Hence, in this work, we aim to explore the capability of LLMs in low resource ASR and Mandarin-English code switching ASR. We also evaluate and compare the recognition performance of LLM-based ASR systems against Whisper model. Extensive experiments demonstrate that LLM-based ASR yields a relative gain of 12.8\% over the Whisper model in low resource ASR while Whisper performs better in Mandarin-English code switching ASR. We hope that this study could shed light on ASR for low resource scenarios. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_00721 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario Song, Zheshu Ma, Ziyang Yang, Yifan Zhuo, Jianheng Chen, Xie Artificial Intelligence Computation and Language Sound Audio and Speech Processing Large Language Models (LLMs) have showcased exceptional performance across diverse NLP tasks, and their integration with speech encoder is rapidly emerging as a dominant trend in the Automatic Speech Recognition (ASR) field. Previous works mainly concentrated on leveraging LLMs for speech recognition in English and Chinese. However, their potential for addressing speech recognition challenges in low resource settings remains underexplored. Hence, in this work, we aim to explore the capability of LLMs in low resource ASR and Mandarin-English code switching ASR. We also evaluate and compare the recognition performance of LLM-based ASR systems against Whisper model. Extensive experiments demonstrate that LLM-based ASR yields a relative gain of 12.8\% over the Whisper model in low resource ASR while Whisper performs better in Mandarin-English code switching ASR. We hope that this study could shed light on ASR for low resource scenarios. |
| title | A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario |
| topic | Artificial Intelligence Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2412.00721 |