PianoVAM: A Multimodal Piano Performance Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Yonghyun, Park, Junhyung, Bae, Joonhyung, Kim, Kirak, Kwon, Taegyun, Lerch, Alexander, Nam, Juhan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915488795197440
author Kim, Yonghyun
Park, Junhyung
Bae, Joonhyung
Kim, Kirak
Kwon, Taegyun
Lerch, Alexander
Nam, Juhan
author_facet Kim, Yonghyun
Park, Junhyung
Bae, Joonhyung
Kim, Kirak
Kwon, Taegyun
Lerch, Alexander
Nam, Juhan
contents The multimodal nature of music performance has driven increasing interest in data beyond the audio domain within the music information retrieval (MIR) community. This paper introduces PianoVAM, a comprehensive piano performance dataset that includes videos, audio, MIDI, hand landmarks, fingering labels, and rich metadata. The dataset was recorded using a Disklavier piano, capturing audio and MIDI from amateur pianists during their daily practice sessions, alongside synchronized top-view videos in realistic and varied performance conditions. Hand landmarks and fingering labels were extracted using a pretrained hand pose estimation model and a semi-automated fingering annotation algorithm. We discuss the challenges encountered during data collection and the alignment process across different modalities. Additionally, we describe our fingering annotation method based on hand landmarks extracted from videos. Finally, we present benchmarking results for both audio-only and audio-visual piano transcription using the PianoVAM dataset and discuss additional potential applications.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08800
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PianoVAM: A Multimodal Piano Performance Dataset
Kim, Yonghyun
Park, Junhyung
Bae, Joonhyung
Kim, Kirak
Kwon, Taegyun
Lerch, Alexander
Nam, Juhan
Sound
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
The multimodal nature of music performance has driven increasing interest in data beyond the audio domain within the music information retrieval (MIR) community. This paper introduces PianoVAM, a comprehensive piano performance dataset that includes videos, audio, MIDI, hand landmarks, fingering labels, and rich metadata. The dataset was recorded using a Disklavier piano, capturing audio and MIDI from amateur pianists during their daily practice sessions, alongside synchronized top-view videos in realistic and varied performance conditions. Hand landmarks and fingering labels were extracted using a pretrained hand pose estimation model and a semi-automated fingering annotation algorithm. We discuss the challenges encountered during data collection and the alignment process across different modalities. Additionally, we describe our fingering annotation method based on hand landmarks extracted from videos. Finally, we present benchmarking results for both audio-only and audio-visual piano transcription using the PianoVAM dataset and discuss additional potential applications.
title PianoVAM: A Multimodal Piano Performance Dataset
topic Sound
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2509.08800