A Mamba-based Network for Semi-supervised Singing Melody Extraction Using Confidence Binary Regularization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Xiaoliang, Dong, Kangjie, Cao, Jingkai, Yu, Shuai, Li, Wei, Yu, Yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910941672636416
author He, Xiaoliang
Dong, Kangjie
Cao, Jingkai
Yu, Shuai
Li, Wei
Yu, Yi
author_facet He, Xiaoliang
Dong, Kangjie
Cao, Jingkai
Yu, Shuai
Li, Wei
Yu, Yi
contents Singing melody extraction (SME) is a key task in the field of music information retrieval. However, existing methods are facing several limitations: firstly, prior models use transformers to capture the contextual dependencies, which requires quadratic computation resulting in low efficiency in the inference stage. Secondly, prior works typically rely on frequencysupervised methods to estimate the fundamental frequency (f0), which ignores that the musical performance is actually based on notes. Thirdly, transformers typically require large amounts of labeled data to achieve optimal performances, but the SME task lacks of sufficient annotated data. To address these issues, in this paper, we propose a mamba-based network, called SpectMamba, for semi-supervised singing melody extraction using confidence binary regularization. In particular, we begin by introducing vision mamba to achieve computational linear complexity. Then, we propose a novel note-f0 decoder that allows the model to better mimic the musical performance. Further, to alleviate the scarcity of the labeled data, we introduce a confidence binary regularization (CBR) module to leverage the unlabeled data by maximizing the probability of the correct classes. The proposed method is evaluated on several public datasets and the conducted experiments demonstrate the effectiveness of our proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08681
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Mamba-based Network for Semi-supervised Singing Melody Extraction Using Confidence Binary Regularization
He, Xiaoliang
Dong, Kangjie
Cao, Jingkai
Yu, Shuai
Li, Wei
Yu, Yi
Sound
Artificial Intelligence
Audio and Speech Processing
Singing melody extraction (SME) is a key task in the field of music information retrieval. However, existing methods are facing several limitations: firstly, prior models use transformers to capture the contextual dependencies, which requires quadratic computation resulting in low efficiency in the inference stage. Secondly, prior works typically rely on frequencysupervised methods to estimate the fundamental frequency (f0), which ignores that the musical performance is actually based on notes. Thirdly, transformers typically require large amounts of labeled data to achieve optimal performances, but the SME task lacks of sufficient annotated data. To address these issues, in this paper, we propose a mamba-based network, called SpectMamba, for semi-supervised singing melody extraction using confidence binary regularization. In particular, we begin by introducing vision mamba to achieve computational linear complexity. Then, we propose a novel note-f0 decoder that allows the model to better mimic the musical performance. Further, to alleviate the scarcity of the labeled data, we introduce a confidence binary regularization (CBR) module to leverage the unlabeled data by maximizing the probability of the correct classes. The proposed method is evaluated on several public datasets and the conducted experiments demonstrate the effectiveness of our proposed method.
title A Mamba-based Network for Semi-supervised Singing Melody Extraction Using Confidence Binary Regularization
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.08681