Generating Piano Music with Transformers: A Comparative Study of Scale, Data, and Metrics
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912801320075264 |
|---|---|
| author | Lehmkuhl, Jonathan Ilyés-Kun, Ábel Bremes, Nico Özaltan, Cemhan Kaan Muthers, Frederik Yuan, Jiayi |
| author_facet | Lehmkuhl, Jonathan Ilyés-Kun, Ábel Bremes, Nico Özaltan, Cemhan Kaan Muthers, Frederik Yuan, Jiayi |
| contents | Although a variety of transformers have been proposed for symbolic music generation in recent years, there is still little comprehensive study on how specific design choices affect the quality of the generated music. In this work, we systematically compare different datasets, model architectures, model sizes, and training strategies for the task of symbolic piano music generation. To support model development and evaluation, we examine a range of quantitative metrics and analyze how well they correlate with human judgment collected through listening studies. Our best-performing model, a 950M-parameter transformer trained on 80K MIDI files from diverse genres, produces outputs that are often rated as human-composed in a Turing-style listening survey. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_07268 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Generating Piano Music with Transformers: A Comparative Study of Scale, Data, and Metrics Lehmkuhl, Jonathan Ilyés-Kun, Ábel Bremes, Nico Özaltan, Cemhan Kaan Muthers, Frederik Yuan, Jiayi Sound Audio and Speech Processing Although a variety of transformers have been proposed for symbolic music generation in recent years, there is still little comprehensive study on how specific design choices affect the quality of the generated music. In this work, we systematically compare different datasets, model architectures, model sizes, and training strategies for the task of symbolic piano music generation. To support model development and evaluation, we examine a range of quantitative metrics and analyze how well they correlate with human judgment collected through listening studies. Our best-performing model, a 950M-parameter transformer trained on 80K MIDI files from diverse genres, produces outputs that are often rated as human-composed in a Turing-style listening survey. |
| title | Generating Piano Music with Transformers: A Comparative Study of Scale, Data, and Metrics |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2511.07268 |