Efficient Multilingual ASR Finetuning via LoRA Language Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jiahong, Shao, Yiwen, Zhuo, Jianheng, Li, Chenda, Tang, Liliang, Yu, Dong, Qian, Yanmin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911025743265792
author Li, Jiahong
Shao, Yiwen
Zhuo, Jianheng
Li, Chenda
Tang, Liliang
Yu, Dong
Qian, Yanmin
author_facet Li, Jiahong
Shao, Yiwen
Zhuo, Jianheng
Li, Chenda
Tang, Liliang
Yu, Dong
Qian, Yanmin
contents Recent advancements in deep learning have significantly enhanced multilingual automatic speech recognition (ASR) due to the development of advanced model architectures and available large-scale multilingual datasets. Despite that, multilingual ASR still suffers from the curse of multilinguality in that different languages tend to interfere with each other, making it difficult for the ASR model to identify multiple languages effectively while sharing model capacity across them. This paper proposes an efficient finetuning framework for customized multilingual ASR via prepared LoRA language experts based on Whisper. Through LoRA expert fusion or knowledge distillation, our approach achieves better recognition performance on target languages than standard fine-tuning methods. Experimental results demonstrate that the proposed models yield approximately 10\% and 15\% relative performance gains in language-aware and language-agnostic scenarios, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21555
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Multilingual ASR Finetuning via LoRA Language Experts
Li, Jiahong
Shao, Yiwen
Zhuo, Jianheng
Li, Chenda
Tang, Liliang
Yu, Dong
Qian, Yanmin
Computation and Language
Sound
Audio and Speech Processing
Recent advancements in deep learning have significantly enhanced multilingual automatic speech recognition (ASR) due to the development of advanced model architectures and available large-scale multilingual datasets. Despite that, multilingual ASR still suffers from the curse of multilinguality in that different languages tend to interfere with each other, making it difficult for the ASR model to identify multiple languages effectively while sharing model capacity across them. This paper proposes an efficient finetuning framework for customized multilingual ASR via prepared LoRA language experts based on Whisper. Through LoRA expert fusion or knowledge distillation, our approach achieves better recognition performance on target languages than standard fine-tuning methods. Experimental results demonstrate that the proposed models yield approximately 10\% and 15\% relative performance gains in language-aware and language-agnostic scenarios, respectively.
title Efficient Multilingual ASR Finetuning via LoRA Language Experts
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.21555