VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yukun, Wang, Tianrui, Mu, Zhaoxi, Yang, Xinyu, Chng, EngSiong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913094211469312
author Chen, Yukun
Wang, Tianrui
Mu, Zhaoxi
Yang, Xinyu
Chng, EngSiong
author_facet Chen, Yukun
Wang, Tianrui
Mu, Zhaoxi
Yang, Xinyu
Chng, EngSiong
contents High-quality singing annotations are fundamental to modern Singing Voice Synthesis (SVS) systems. However, obtaining these annotations at scale through manual labeling is unrealistic due to the substantial labor and musical expertise required, making automatic annotation highly necessary. Despite their utility, current automatic transcription systems face significant challenges: they often rely on complex multi-stage pipelines, struggle to recover text-note alignments, and exhibit poor generalization to out-of-distribution (OOD) singing data. To alleviate these issues, we present VocalParse, a unified singing voice transcription (SVT) model built upon a Large Audio Language Model (LALM). Specifically, our novel contribution is to introduce an interleaved prompting formulation that jointly models lyrics, melody, and word-note correspondence, yielding a generated sequence that directly maps to a structured musical score. Furthermore, we propose a Chain-of-Thought (CoT) style prompting strategy, which decodes lyrics first as a semantic scaffold, significantly mitigating the context disruption problem while preserving the structural benefits of interleaved generation. Experiments demonstrate that VocalParse achieves state-of-the-art SVT performance on multiple singing datasets. The source code and checkpoint are available at https://github.com/pymaster17/VocalParse.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04613
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models
Chen, Yukun
Wang, Tianrui
Mu, Zhaoxi
Yang, Xinyu
Chng, EngSiong
Sound
Artificial Intelligence
High-quality singing annotations are fundamental to modern Singing Voice Synthesis (SVS) systems. However, obtaining these annotations at scale through manual labeling is unrealistic due to the substantial labor and musical expertise required, making automatic annotation highly necessary. Despite their utility, current automatic transcription systems face significant challenges: they often rely on complex multi-stage pipelines, struggle to recover text-note alignments, and exhibit poor generalization to out-of-distribution (OOD) singing data. To alleviate these issues, we present VocalParse, a unified singing voice transcription (SVT) model built upon a Large Audio Language Model (LALM). Specifically, our novel contribution is to introduce an interleaved prompting formulation that jointly models lyrics, melody, and word-note correspondence, yielding a generated sequence that directly maps to a structured musical score. Furthermore, we propose a Chain-of-Thought (CoT) style prompting strategy, which decodes lyrics first as a semantic scaffold, significantly mitigating the context disruption problem while preserving the structural benefits of interleaved generation. Experiments demonstrate that VocalParse achieves state-of-the-art SVT performance on multiple singing datasets. The source code and checkpoint are available at https://github.com/pymaster17/VocalParse.
title VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2605.04613