Moonbeam: A MIDI Foundation Model Using Both Absolute and Relative Music Attributes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Zixun, Dixon, Simon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915296264060928
author Guo, Zixun
Dixon, Simon
author_facet Guo, Zixun
Dixon, Simon
contents Moonbeam is a transformer-based foundation model for symbolic music, pretrained on a large and diverse collection of MIDI data totaling 81.6K hours of music and 18 billion tokens. Moonbeam incorporates music-domain inductive biases by capturing both absolute and relative musical attributes through the introduction of a novel domain-knowledge-inspired tokenization method and Multidimensional Relative Attention (MRA), which captures relative music information without additional trainable parameters. Leveraging the pretrained Moonbeam, we propose 2 finetuning architectures with full anticipatory capabilities, targeting 2 categories of downstream tasks: symbolic music understanding and conditional music generation (including music infilling). Our model outperforms other large-scale pretrained music models in most cases in terms of accuracy and F1 score across 3 downstream music classification tasks on 4 datasets. Moreover, our finetuned conditional music generation model outperforms a strong transformer baseline with a REMI-like tokenizer. We open-source the code, pretrained model, and generated samples on Github.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15559
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Moonbeam: A MIDI Foundation Model Using Both Absolute and Relative Music Attributes
Guo, Zixun
Dixon, Simon
Sound
Artificial Intelligence
Audio and Speech Processing
Moonbeam is a transformer-based foundation model for symbolic music, pretrained on a large and diverse collection of MIDI data totaling 81.6K hours of music and 18 billion tokens. Moonbeam incorporates music-domain inductive biases by capturing both absolute and relative musical attributes through the introduction of a novel domain-knowledge-inspired tokenization method and Multidimensional Relative Attention (MRA), which captures relative music information without additional trainable parameters. Leveraging the pretrained Moonbeam, we propose 2 finetuning architectures with full anticipatory capabilities, targeting 2 categories of downstream tasks: symbolic music understanding and conditional music generation (including music infilling). Our model outperforms other large-scale pretrained music models in most cases in terms of accuracy and F1 score across 3 downstream music classification tasks on 4 datasets. Moreover, our finetuned conditional music generation model outperforms a strong transformer baseline with a REMI-like tokenizer. We open-source the code, pretrained model, and generated samples on Github.
title Moonbeam: A MIDI Foundation Model Using Both Absolute and Relative Music Attributes
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.15559