Is Transfer Learning Necessary for Violin Transcription?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Yueh-Po, Wang, Ting-Kang, Su, Li, Cheung, Vincent K. M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911111962427392
author Peng, Yueh-Po
Wang, Ting-Kang
Su, Li
Cheung, Vincent K. M.
author_facet Peng, Yueh-Po
Wang, Ting-Kang
Su, Li
Cheung, Vincent K. M.
contents Automatic music transcription (AMT) has achieved remarkable progress for instruments such as the piano, largely due to the availability of large-scale, high-quality datasets. In contrast, violin AMT remains underexplored due to limited annotated data. A common approach is to fine-tune pretrained models for other downstream tasks, but the effectiveness of such transfer remains unclear in the presence of timbral and articulatory differences. In this work, we investigate whether training from scratch on a medium-scale violin dataset can match the performance of fine-tuned piano-pretrained models. We adopt a piano transcription architecture without modification and train it on the MOSA dataset, which contains about 30 hours of aligned violin recordings. Our experiments on URMP and Bach10 show that models trained from scratch achieved competitive or even superior performance compared to fine-tuned counterparts. These findings suggest that strong violin AMT is possible without relying on pretrained piano representations, highlighting the importance of instrument-specific data collection and augmentation strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13516
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Is Transfer Learning Necessary for Violin Transcription?
Peng, Yueh-Po
Wang, Ting-Kang
Su, Li
Cheung, Vincent K. M.
Sound
Audio and Speech Processing
Automatic music transcription (AMT) has achieved remarkable progress for instruments such as the piano, largely due to the availability of large-scale, high-quality datasets. In contrast, violin AMT remains underexplored due to limited annotated data. A common approach is to fine-tune pretrained models for other downstream tasks, but the effectiveness of such transfer remains unclear in the presence of timbral and articulatory differences. In this work, we investigate whether training from scratch on a medium-scale violin dataset can match the performance of fine-tuned piano-pretrained models. We adopt a piano transcription architecture without modification and train it on the MOSA dataset, which contains about 30 hours of aligned violin recordings. Our experiments on URMP and Bach10 show that models trained from scratch achieved competitive or even superior performance compared to fine-tuned counterparts. These findings suggest that strong violin AMT is possible without relying on pretrained piano representations, highlighting the importance of instrument-specific data collection and augmentation strategies.
title Is Transfer Learning Necessary for Violin Transcription?
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2508.13516