Exploring the potential and limitations of Model Merging for Multi-Domain Adaptation in ASR
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917316226187264 |
|---|---|
| author | Carvalho, Carlos Teixeira, Francisco Rolland, Thomas Abad, Alberto |
| author_facet | Carvalho, Carlos Teixeira, Francisco Rolland, Thomas Abad, Alberto |
| contents | Model merging is a scalable alternative to multi-task training that combines the capabilities of multiple specialised models into a single model. This is particularly attractive for large speech foundation models, which are typically adapted through domain-specific fine-tuning, resulting in multiple customised checkpoints, for which repeating full fine-tuning when new data becomes available is computationally prohibitive. In this work, we study model merging for multi-domain ASR and benchmark 11 merging algorithms for 10 European Portuguese domains, evaluating in-domain accuracy, robustness under distribution shift, as well as English and multilingual performance. We further propose BoostedTSV-M, a new merging algorithm based on TSV-M that mitigates rank collapse via singular-value boosting and improves numerical stability. Overall, our approach outperforms full fine-tuning on European Portuguese while preserving out-of-distribution generalisation in a single model. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_05354 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Exploring the potential and limitations of Model Merging for Multi-Domain Adaptation in ASR Carvalho, Carlos Teixeira, Francisco Rolland, Thomas Abad, Alberto Computation and Language Audio and Speech Processing Model merging is a scalable alternative to multi-task training that combines the capabilities of multiple specialised models into a single model. This is particularly attractive for large speech foundation models, which are typically adapted through domain-specific fine-tuning, resulting in multiple customised checkpoints, for which repeating full fine-tuning when new data becomes available is computationally prohibitive. In this work, we study model merging for multi-domain ASR and benchmark 11 merging algorithms for 10 European Portuguese domains, evaluating in-domain accuracy, robustness under distribution shift, as well as English and multilingual performance. We further propose BoostedTSV-M, a new merging algorithm based on TSV-M that mitigates rank collapse via singular-value boosting and improves numerical stability. Overall, our approach outperforms full fine-tuning on European Portuguese while preserving out-of-distribution generalisation in a single model. |
| title | Exploring the potential and limitations of Model Merging for Multi-Domain Adaptation in ASR |
| topic | Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2603.05354 |