Robust fine-tuning of speech recognition models via model merging: application to disordered speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ducorroy, Alexandre, Riad, Rachid
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913859959259136
author Ducorroy, Alexandre
Riad, Rachid
author_facet Ducorroy, Alexandre
Riad, Rachid
contents Automatic Speech Recognition (ASR) has advanced with Speech Foundation Models (SFMs), yet performance degrades on dysarthric speech due to variability and limited data. This study as part of the submission to the Speech Accessibility challenge, explored model merging to improve ASR generalization using Whisper as the base SFM. We compared fine-tuning with single-trajectory merging, combining models from one fine-tuning path, and multi-run merging, merging independently trained models. Our best multi-run merging approach achieved a 12% relative decrease of WER over classic fine-tuning, and a 16.2% relative decrease on long-form audios, a major loss contributor in dysarthric ASR. Merging more and more models led to continuous gains, remained effective in low-data regimes, and generalized across model architectures. These results highlight model merging as an easily replicable adaptation method that consistently improves ASR without additional inference cost or hyperparameter tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20477
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robust fine-tuning of speech recognition models via model merging: application to disordered speech
Ducorroy, Alexandre
Riad, Rachid
Audio and Speech Processing
Sound
Automatic Speech Recognition (ASR) has advanced with Speech Foundation Models (SFMs), yet performance degrades on dysarthric speech due to variability and limited data. This study as part of the submission to the Speech Accessibility challenge, explored model merging to improve ASR generalization using Whisper as the base SFM. We compared fine-tuning with single-trajectory merging, combining models from one fine-tuning path, and multi-run merging, merging independently trained models. Our best multi-run merging approach achieved a 12% relative decrease of WER over classic fine-tuning, and a 16.2% relative decrease on long-form audios, a major loss contributor in dysarthric ASR. Merging more and more models led to continuous gains, remained effective in low-data regimes, and generalized across model architectures. These results highlight model merging as an easily replicable adaptation method that consistently improves ASR without additional inference cost or hyperparameter tuning.
title Robust fine-tuning of speech recognition models via model merging: application to disordered speech
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2505.20477