MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music Composition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pasquier, Philippe, Ens, Jeff, Fradet, Nathan, Triana, Paul, Rizzotti, Davide, Rolland, Jean-Baptiste, Safi, Maryam
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910813402431488
author Pasquier, Philippe
Ens, Jeff
Fradet, Nathan
Triana, Paul
Rizzotti, Davide
Rolland, Jean-Baptiste
Safi, Maryam
author_facet Pasquier, Philippe
Ens, Jeff
Fradet, Nathan
Triana, Paul
Rizzotti, Davide
Rolland, Jean-Baptiste
Safi, Maryam
contents We present and release MIDI-GPT, a generative system based on the Transformer architecture that is designed for computer-assisted music composition workflows. MIDI-GPT supports the infilling of musical material at the track and bar level, and can condition generation on attributes including: instrument type, musical style, note density, polyphony level, and note duration. In order to integrate these features, we employ an alternative representation for musical material, creating a time-ordered sequence of musical events for each track and concatenating several tracks into a single sequence, rather than using a single time-ordered sequence where the musical events corresponding to different tracks are interleaved. We also propose a variation of our representation allowing for expressiveness. We present experimental results that demonstrate that MIDI-GPT is able to consistently avoid duplicating the musical material it was trained on, generate music that is stylistically similar to the training dataset, and that attribute controls allow enforcing various constraints on the generated material. We also outline several real-world applications of MIDI-GPT, including collaborations with industry partners that explore the integration and evaluation of MIDI-GPT into commercial products, as well as several artistic works produced using it.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17011
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music Composition
Pasquier, Philippe
Ens, Jeff
Fradet, Nathan
Triana, Paul
Rizzotti, Davide
Rolland, Jean-Baptiste
Safi, Maryam
Sound
Machine Learning
Multimedia
Audio and Speech Processing
We present and release MIDI-GPT, a generative system based on the Transformer architecture that is designed for computer-assisted music composition workflows. MIDI-GPT supports the infilling of musical material at the track and bar level, and can condition generation on attributes including: instrument type, musical style, note density, polyphony level, and note duration. In order to integrate these features, we employ an alternative representation for musical material, creating a time-ordered sequence of musical events for each track and concatenating several tracks into a single sequence, rather than using a single time-ordered sequence where the musical events corresponding to different tracks are interleaved. We also propose a variation of our representation allowing for expressiveness. We present experimental results that demonstrate that MIDI-GPT is able to consistently avoid duplicating the musical material it was trained on, generate music that is stylistically similar to the training dataset, and that attribute controls allow enforcing various constraints on the generated material. We also outline several real-world applications of MIDI-GPT, including collaborations with industry partners that explore the integration and evaluation of MIDI-GPT into commercial products, as well as several artistic works produced using it.
title MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music Composition
topic Sound
Machine Learning
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2501.17011