Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866913635982376960 |
|---|---|
| author | Zhao, Jiahui Shi, Hao Cui, Chenrui Wang, Tianrui Liu, Hexin Ni, Zhaoheng Ye, Lingxuan Wang, Longbiao |
| author_facet | Zhao, Jiahui Shi, Hao Cui, Chenrui Wang, Tianrui Liu, Hexin Ni, Zhaoheng Ye, Lingxuan Wang, Longbiao |
| contents | Code-switching (CS) automatic speech recognition (ASR) faces challenges due to the language confusion resulting from accents, auditory similarity, and seamless language switches. Adaptation on the pre-trained multi-lingual model has shown promising performance for CS-ASR. In this paper, we adapt Whisper, which is a large-scale multilingual pre-trained speech recognition model, to CS from both encoder and decoder parts. First, we propose an encoder refiner to enhance the encoder's capacity of intra-sentence swithching. Second, we propose using two sets of language-aware adapters with different language prompt embeddings to achieve language-specific decoding information in each decoder layer. Then, a fusion module is added to fuse the language-aware decoding. The experimental results using the SEAME dataset show that, compared with the baseline model, the proposed approach achieves a relative MER reduction of 4.1% and 7.2% on the dev_man and dev_sge test sets, respectively, surpassing state-of-the-art methods. Through experiments, we found that the proposed method significantly improves the performance on non-native language in CS speech, indicating that our approach enables Whisper to better distinguish between the two languages. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_16507 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding Zhao, Jiahui Shi, Hao Cui, Chenrui Wang, Tianrui Liu, Hexin Ni, Zhaoheng Ye, Lingxuan Wang, Longbiao Computation and Language Sound Audio and Speech Processing Code-switching (CS) automatic speech recognition (ASR) faces challenges due to the language confusion resulting from accents, auditory similarity, and seamless language switches. Adaptation on the pre-trained multi-lingual model has shown promising performance for CS-ASR. In this paper, we adapt Whisper, which is a large-scale multilingual pre-trained speech recognition model, to CS from both encoder and decoder parts. First, we propose an encoder refiner to enhance the encoder's capacity of intra-sentence swithching. Second, we propose using two sets of language-aware adapters with different language prompt embeddings to achieve language-specific decoding information in each decoder layer. Then, a fusion module is added to fuse the language-aware decoding. The experimental results using the SEAME dataset show that, compared with the baseline model, the proposed approach achieves a relative MER reduction of 4.1% and 7.2% on the dev_man and dev_sge test sets, respectively, surpassing state-of-the-art methods. Through experiments, we found that the proposed method significantly improves the performance on non-native language in CS speech, indicating that our approach enables Whisper to better distinguish between the two languages. |
| title | Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2412.16507 |