SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917708086378496 |
|---|---|
| author | Zhao, Qiuming Sun, Guangzhi Zhang, Chao Xu, Mingxing Zheng, Thomas Fang |
| author_facet | Zhao, Qiuming Sun, Guangzhi Zhang, Chao Xu, Mingxing Zheng, Thomas Fang |
| contents | Mixture-of-experts (MoE) models have achieved excellent results in many tasks. However, conventional MoE models are often very large, making them challenging to deploy on resource-constrained edge devices. In this paper, we propose a novel speaker adaptive mixture of LoRA experts (SAML) approach, which uses low-rank adaptation (LoRA) modules as experts to reduce the number of trainable parameters in MoE. Specifically, SAML is applied to the quantised and personalised end-to-end automatic speech recognition models, which combines test-time speaker adaptation to improve the performance of heavily compressed models in speaker-specific scenarios. Experiments have been performed on the LibriSpeech and the TED-LIUM 3 corpora. Remarkably, with a 7x reduction in model size, 29.1% and 31.1% relative word error rate reductions were achieved on the quantised Whisper model and Conformer-based attention-based encoder-decoder ASR model respectively, comparing to the original full precision models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_19706 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR Zhao, Qiuming Sun, Guangzhi Zhang, Chao Xu, Mingxing Zheng, Thomas Fang Sound Audio and Speech Processing Mixture-of-experts (MoE) models have achieved excellent results in many tasks. However, conventional MoE models are often very large, making them challenging to deploy on resource-constrained edge devices. In this paper, we propose a novel speaker adaptive mixture of LoRA experts (SAML) approach, which uses low-rank adaptation (LoRA) modules as experts to reduce the number of trainable parameters in MoE. Specifically, SAML is applied to the quantised and personalised end-to-end automatic speech recognition models, which combines test-time speaker adaptation to improve the performance of heavily compressed models in speaker-specific scenarios. Experiments have been performed on the LibriSpeech and the TED-LIUM 3 corpora. Remarkably, with a 7x reduction in model size, 29.1% and 31.1% relative word error rate reductions were achieved on the quantised Whisper model and Conformer-based attention-based encoder-decoder ASR model respectively, comparing to the original full precision models. |
| title | SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2406.19706 |