Annotation-free Automatic Music Transcription with Scalable Synthetic Data and Adversarial Domain Confusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sato, Gakusei, Akama, Taketo
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917711008759808
author Sato, Gakusei
Akama, Taketo
author_facet Sato, Gakusei
Akama, Taketo
contents Automatic Music Transcription (AMT) is a vital technology in the field of music information processing. Despite recent enhancements in performance due to machine learning techniques, current methods typically attain high accuracy in domains where abundant annotated data is available. Addressing domains with low or no resources continues to be an unresolved challenge. To tackle this issue, we propose a transcription model that does not require any MIDI-audio paired data through the utilization of scalable synthetic audio for pre-training and adversarial domain confusion using unannotated real audio. In experiments, we evaluate methods under the real-world application scenario where training datasets do not include the MIDI annotation of audio in the target data domain. Our proposed method achieved competitive performance relative to established baseline methods, despite not utilizing any real datasets of paired MIDI-audio. Additionally, ablation studies have provided insights into the scalability of this approach and the forthcoming challenges in the field of AMT research.
format Preprint
id arxiv_https___arxiv_org_abs_2312_10402
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Annotation-free Automatic Music Transcription with Scalable Synthetic Data and Adversarial Domain Confusion
Sato, Gakusei
Akama, Taketo
Sound
Artificial Intelligence
Audio and Speech Processing
Automatic Music Transcription (AMT) is a vital technology in the field of music information processing. Despite recent enhancements in performance due to machine learning techniques, current methods typically attain high accuracy in domains where abundant annotated data is available. Addressing domains with low or no resources continues to be an unresolved challenge. To tackle this issue, we propose a transcription model that does not require any MIDI-audio paired data through the utilization of scalable synthetic audio for pre-training and adversarial domain confusion using unannotated real audio. In experiments, we evaluate methods under the real-world application scenario where training datasets do not include the MIDI annotation of audio in the target data domain. Our proposed method achieved competitive performance relative to established baseline methods, despite not utilizing any real datasets of paired MIDI-audio. Additionally, ablation studies have provided insights into the scalability of this approach and the forthcoming challenges in the field of AMT research.
title Annotation-free Automatic Music Transcription with Scalable Synthetic Data and Adversarial Domain Confusion
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2312.10402