Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bradshaw, Louis, Fan, Honglu, Spangher, Alexander, Biderman, Stella, Colton, Simon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916817440604160
author Bradshaw, Louis
Fan, Honglu
Spangher, Alexander
Biderman, Stella
Colton, Simon
author_facet Bradshaw, Louis
Fan, Honglu
Spangher, Alexander
Biderman, Stella
Colton, Simon
contents We study the capabilities of generative autoregressive transformer models trained on large amounts of symbolic solo-piano transcriptions. After first pretraining on approximately 60,000 hours of music, we use a comparatively smaller, high-quality subset, to finetune models to produce musical continuations, perform symbolic classification tasks, and produce general-purpose contrastive MIDI embeddings by adapting the SimCLR framework to symbolic music. When evaluating piano continuation coherence, our generative model outperforms leading symbolic generation techniques and remains competitive with proprietary audio generation models. On MIR classification benchmarks, frozen representations from our contrastive model achieve state-of-the-art results in linear probe experiments, while direct finetuning demonstrates the generalizability of pretrained representations, often requiring only a few hundred labeled examples to specialize to downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23869
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
Bradshaw, Louis
Fan, Honglu
Spangher, Alexander
Biderman, Stella
Colton, Simon
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
We study the capabilities of generative autoregressive transformer models trained on large amounts of symbolic solo-piano transcriptions. After first pretraining on approximately 60,000 hours of music, we use a comparatively smaller, high-quality subset, to finetune models to produce musical continuations, perform symbolic classification tasks, and produce general-purpose contrastive MIDI embeddings by adapting the SimCLR framework to symbolic music. When evaluating piano continuation coherence, our generative model outperforms leading symbolic generation techniques and remains competitive with proprietary audio generation models. On MIR classification benchmarks, frozen representations from our contrastive model achieve state-of-the-art results in linear probe experiments, while direct finetuning demonstrates the generalizability of pretrained representations, often requiring only a few hundred labeled examples to specialize to downstream tasks.
title Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2506.23869