On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: HU, Shujie, Xie, Xurong, Geng, Mengzhe, Deng, Jiajun, Wang, Huimeng, Li, Guinan, Deng, Chengxi, Wang, Tianzi, Cui, Mingyu, Meng, Helen, Liu, Xunying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909625923665920
author HU, Shujie
Xie, Xurong
Geng, Mengzhe
Deng, Jiajun
Wang, Huimeng
Li, Guinan
Deng, Chengxi
Wang, Tianzi
Cui, Mingyu
Meng, Helen
Liu, Xunying
author_facet HU, Shujie
Xie, Xurong
Geng, Mengzhe
Deng, Jiajun
Wang, Huimeng
Li, Guinan
Deng, Chengxi
Wang, Tianzi
Cui, Mingyu
Meng, Helen
Liu, Xunying
contents This paper proposes a novel MoE-based speaker adaptation framework for foundation models based dysarthric speech recognition. This approach enables zero-shot adaptation and real-time processing while incorporating domain knowledge. Speech impairment severity and gender conditioned adapter experts are dynamically combined using on-the-fly predicted speaker-dependent routing parameters. KL-divergence is used to further enforce diversity among experts and their generalization to unseen speakers. Experimental results on the UASpeech corpus suggest that on-the-fly MoE-based adaptation produces statistically significant WER reductions of up to 1.34% absolute (6.36% relative) over the unadapted baseline HuBERT/WavLM models. Consistent WER reductions of up to 2.55% absolute (11.44% relative) and RTF speedups of up to 7 times are obtained over batch-mode adaptation across varying speaker-level data quantities. The lowest published WER of 16.35% (46.77% on very low intelligibility) is obtained.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22072
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
HU, Shujie
Xie, Xurong
Geng, Mengzhe
Deng, Jiajun
Wang, Huimeng
Li, Guinan
Deng, Chengxi
Wang, Tianzi
Cui, Mingyu
Meng, Helen
Liu, Xunying
Sound
Audio and Speech Processing
This paper proposes a novel MoE-based speaker adaptation framework for foundation models based dysarthric speech recognition. This approach enables zero-shot adaptation and real-time processing while incorporating domain knowledge. Speech impairment severity and gender conditioned adapter experts are dynamically combined using on-the-fly predicted speaker-dependent routing parameters. KL-divergence is used to further enforce diversity among experts and their generalization to unseen speakers. Experimental results on the UASpeech corpus suggest that on-the-fly MoE-based adaptation produces statistically significant WER reductions of up to 1.34% absolute (6.36% relative) over the unadapted baseline HuBERT/WavLM models. Consistent WER reductions of up to 2.55% absolute (11.44% relative) and RTF speedups of up to 7 times are obtained over batch-mode adaptation across varying speaker-level data quantities. The lowest published WER of 16.35% (46.77% on very low intelligibility) is obtained.
title On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.22072