Mixture of LoRA Experts for Low-Resourced Multi-Accent Automatic Speech Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bagat, Raphaël, Illina, Irina, Vincent, Emmanuel
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908572553576448
author Bagat, Raphaël
Illina, Irina
Vincent, Emmanuel
author_facet Bagat, Raphaël
Illina, Irina
Vincent, Emmanuel
contents We aim to improve the robustness of Automatic Speech Recognition (ASR) systems against non-native speech, particularly in low-resourced multi-accent settings. We introduce Mixture of Accent-Specific LoRAs (MAS-LoRA), a fine-tuning method that leverages a mixture of Low-Rank Adaptation (LoRA) experts, each specialized in a specific accent. This method can be used when the accent is known or unknown at inference time, without the need to fine-tune the model again. Our experiments, conducted using Whisper on the L2-ARCTIC corpus, demonstrate significant improvements in Word Error Rate compared to regular LoRA and full fine-tuning when the accent is unknown. When the accent is known, the results further improve. Furthermore, MAS-LoRA shows less catastrophic forgetting than the other fine-tuning methods. To the best of our knowledge, this is the first use of a mixture of LoRA experts for non-native multi-accent ASR.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20006
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mixture of LoRA Experts for Low-Resourced Multi-Accent Automatic Speech Recognition
Bagat, Raphaël
Illina, Irina
Vincent, Emmanuel
Computation and Language
We aim to improve the robustness of Automatic Speech Recognition (ASR) systems against non-native speech, particularly in low-resourced multi-accent settings. We introduce Mixture of Accent-Specific LoRAs (MAS-LoRA), a fine-tuning method that leverages a mixture of Low-Rank Adaptation (LoRA) experts, each specialized in a specific accent. This method can be used when the accent is known or unknown at inference time, without the need to fine-tune the model again. Our experiments, conducted using Whisper on the L2-ARCTIC corpus, demonstrate significant improvements in Word Error Rate compared to regular LoRA and full fine-tuning when the accent is unknown. When the accent is known, the results further improve. Furthermore, MAS-LoRA shows less catastrophic forgetting than the other fine-tuning methods. To the best of our knowledge, this is the first use of a mixture of LoRA experts for non-native multi-accent ASR.
title Mixture of LoRA Experts for Low-Resourced Multi-Accent Automatic Speech Recognition
topic Computation and Language
url https://arxiv.org/abs/2505.20006