A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Song, Zheshu, Ma, Ziyang, Yang, Yifan, Zhuo, Jianheng, Chen, Xie
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909414849511424
author Song, Zheshu
Ma, Ziyang
Yang, Yifan
Zhuo, Jianheng
Chen, Xie
author_facet Song, Zheshu
Ma, Ziyang
Yang, Yifan
Zhuo, Jianheng
Chen, Xie
contents Large Language Models (LLMs) have showcased exceptional performance across diverse NLP tasks, and their integration with speech encoder is rapidly emerging as a dominant trend in the Automatic Speech Recognition (ASR) field. Previous works mainly concentrated on leveraging LLMs for speech recognition in English and Chinese. However, their potential for addressing speech recognition challenges in low resource settings remains underexplored. Hence, in this work, we aim to explore the capability of LLMs in low resource ASR and Mandarin-English code switching ASR. We also evaluate and compare the recognition performance of LLM-based ASR systems against Whisper model. Extensive experiments demonstrate that LLM-based ASR yields a relative gain of 12.8\% over the Whisper model in low resource ASR while Whisper performs better in Mandarin-English code switching ASR. We hope that this study could shed light on ASR for low resource scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00721
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
Song, Zheshu
Ma, Ziyang
Yang, Yifan
Zhuo, Jianheng
Chen, Xie
Artificial Intelligence
Computation and Language
Sound
Audio and Speech Processing
Large Language Models (LLMs) have showcased exceptional performance across diverse NLP tasks, and their integration with speech encoder is rapidly emerging as a dominant trend in the Automatic Speech Recognition (ASR) field. Previous works mainly concentrated on leveraging LLMs for speech recognition in English and Chinese. However, their potential for addressing speech recognition challenges in low resource settings remains underexplored. Hence, in this work, we aim to explore the capability of LLMs in low resource ASR and Mandarin-English code switching ASR. We also evaluate and compare the recognition performance of LLM-based ASR systems against Whisper model. Extensive experiments demonstrate that LLM-based ASR yields a relative gain of 12.8\% over the Whisper model in low resource ASR while Whisper performs better in Mandarin-English code switching ASR. We hope that this study could shed light on ASR for low resource scenarios.
title A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
topic Artificial Intelligence
Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.00721