Generating Piano Music with Transformers: A Comparative Study of Scale, Data, and Metrics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lehmkuhl, Jonathan, Ilyés-Kun, Ábel, Bremes, Nico, Özaltan, Cemhan Kaan, Muthers, Frederik, Yuan, Jiayi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912801320075264
author Lehmkuhl, Jonathan
Ilyés-Kun, Ábel
Bremes, Nico
Özaltan, Cemhan Kaan
Muthers, Frederik
Yuan, Jiayi
author_facet Lehmkuhl, Jonathan
Ilyés-Kun, Ábel
Bremes, Nico
Özaltan, Cemhan Kaan
Muthers, Frederik
Yuan, Jiayi
contents Although a variety of transformers have been proposed for symbolic music generation in recent years, there is still little comprehensive study on how specific design choices affect the quality of the generated music. In this work, we systematically compare different datasets, model architectures, model sizes, and training strategies for the task of symbolic piano music generation. To support model development and evaluation, we examine a range of quantitative metrics and analyze how well they correlate with human judgment collected through listening studies. Our best-performing model, a 950M-parameter transformer trained on 80K MIDI files from diverse genres, produces outputs that are often rated as human-composed in a Turing-style listening survey.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07268
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generating Piano Music with Transformers: A Comparative Study of Scale, Data, and Metrics
Lehmkuhl, Jonathan
Ilyés-Kun, Ábel
Bremes, Nico
Özaltan, Cemhan Kaan
Muthers, Frederik
Yuan, Jiayi
Sound
Audio and Speech Processing
Although a variety of transformers have been proposed for symbolic music generation in recent years, there is still little comprehensive study on how specific design choices affect the quality of the generated music. In this work, we systematically compare different datasets, model architectures, model sizes, and training strategies for the task of symbolic piano music generation. To support model development and evaluation, we examine a range of quantitative metrics and analyze how well they correlate with human judgment collected through listening studies. Our best-performing model, a 950M-parameter transformer trained on 80K MIDI files from diverse genres, produces outputs that are often rated as human-composed in a Turing-style listening survey.
title Generating Piano Music with Transformers: A Comparative Study of Scale, Data, and Metrics
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2511.07268