Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhao, Jiahui, Shi, Hao, Cui, Chenrui, Wang, Tianrui, Liu, Hexin, Ni, Zhaoheng, Ye, Lingxuan, Wang, Longbiao
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913635982376960
author Zhao, Jiahui
Shi, Hao
Cui, Chenrui
Wang, Tianrui
Liu, Hexin
Ni, Zhaoheng
Ye, Lingxuan
Wang, Longbiao
author_facet Zhao, Jiahui
Shi, Hao
Cui, Chenrui
Wang, Tianrui
Liu, Hexin
Ni, Zhaoheng
Ye, Lingxuan
Wang, Longbiao
contents Code-switching (CS) automatic speech recognition (ASR) faces challenges due to the language confusion resulting from accents, auditory similarity, and seamless language switches. Adaptation on the pre-trained multi-lingual model has shown promising performance for CS-ASR. In this paper, we adapt Whisper, which is a large-scale multilingual pre-trained speech recognition model, to CS from both encoder and decoder parts. First, we propose an encoder refiner to enhance the encoder's capacity of intra-sentence swithching. Second, we propose using two sets of language-aware adapters with different language prompt embeddings to achieve language-specific decoding information in each decoder layer. Then, a fusion module is added to fuse the language-aware decoding. The experimental results using the SEAME dataset show that, compared with the baseline model, the proposed approach achieves a relative MER reduction of 4.1% and 7.2% on the dev_man and dev_sge test sets, respectively, surpassing state-of-the-art methods. Through experiments, we found that the proposed method significantly improves the performance on non-native language in CS speech, indicating that our approach enables Whisper to better distinguish between the two languages.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16507
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
Zhao, Jiahui
Shi, Hao
Cui, Chenrui
Wang, Tianrui
Liu, Hexin
Ni, Zhaoheng
Ye, Lingxuan
Wang, Longbiao
Computation and Language
Sound
Audio and Speech Processing
Code-switching (CS) automatic speech recognition (ASR) faces challenges due to the language confusion resulting from accents, auditory similarity, and seamless language switches. Adaptation on the pre-trained multi-lingual model has shown promising performance for CS-ASR. In this paper, we adapt Whisper, which is a large-scale multilingual pre-trained speech recognition model, to CS from both encoder and decoder parts. First, we propose an encoder refiner to enhance the encoder's capacity of intra-sentence swithching. Second, we propose using two sets of language-aware adapters with different language prompt embeddings to achieve language-specific decoding information in each decoder layer. Then, a fusion module is added to fuse the language-aware decoding. The experimental results using the SEAME dataset show that, compared with the baseline model, the proposed approach achieves a relative MER reduction of 4.1% and 7.2% on the dev_man and dev_sge test sets, respectively, surpassing state-of-the-art methods. Through experiments, we found that the proposed method significantly improves the performance on non-native language in CS speech, indicating that our approach enables Whisper to better distinguish between the two languages.
title Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.16507