Marco-ASR: A Principled and Metric-Driven Framework for Fine-Tuning Large-Scale ASR Models for Domain Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ni, Xuanfan, Yang, Fei, Tian, Fengping, Li, Qingjuan, Lyu, Chenyang, Du, Yichao, Wang, Longyue, Luo, Weihua, Zhang, Kaifu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911340455526400
author Ni, Xuanfan
Yang, Fei
Tian, Fengping
Li, Qingjuan
Lyu, Chenyang
Du, Yichao
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
author_facet Ni, Xuanfan
Yang, Fei
Tian, Fengping
Li, Qingjuan
Lyu, Chenyang
Du, Yichao
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
contents Automatic Speech Recognition (ASR) models have achieved remarkable accuracy in general settings, yet their performance often degrades in domain-specific applications due to data mismatch and linguistic variability. This challenge is amplified for modern Large Language Model (LLM)-based ASR systems, whose massive scale and complex training dynamics make effective fine-tuning non-trivial. To address this gap, this paper proposes a principled and metric-driven fine-tuning framework for adapting both traditional and LLM-based ASR models to specialized domains. The framework emphasizes learning rate optimization based on performance metrics, combined with domain-specific data transformation and augmentation. We empirically evaluate our framework on state-of-the-art models, including Whisper, Whisper-Turbo, and Qwen2-Audio, across multi-domain, multilingual, and multi-length datasets. Our results not only validate the proposed framework but also establish practical protocols for improving domain-specific ASR performance while preventing overfitting.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22165
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Marco-ASR: A Principled and Metric-Driven Framework for Fine-Tuning Large-Scale ASR Models for Domain Adaptation
Ni, Xuanfan
Yang, Fei
Tian, Fengping
Li, Qingjuan
Lyu, Chenyang
Du, Yichao
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
Sound
Automatic Speech Recognition (ASR) models have achieved remarkable accuracy in general settings, yet their performance often degrades in domain-specific applications due to data mismatch and linguistic variability. This challenge is amplified for modern Large Language Model (LLM)-based ASR systems, whose massive scale and complex training dynamics make effective fine-tuning non-trivial. To address this gap, this paper proposes a principled and metric-driven fine-tuning framework for adapting both traditional and LLM-based ASR models to specialized domains. The framework emphasizes learning rate optimization based on performance metrics, combined with domain-specific data transformation and augmentation. We empirically evaluate our framework on state-of-the-art models, including Whisper, Whisper-Turbo, and Qwen2-Audio, across multi-domain, multilingual, and multi-length datasets. Our results not only validate the proposed framework but also establish practical protocols for improving domain-specific ASR performance while preventing overfitting.
title Marco-ASR: A Principled and Metric-Driven Framework for Fine-Tuning Large-Scale ASR Models for Domain Adaptation
topic Sound
url https://arxiv.org/abs/2512.22165