Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Roy, Abhinaba, Puri, Geeta, Herremans, Dorien
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910951910932480
author Roy, Abhinaba
Puri, Geeta
Herremans, Dorien
author_facet Roy, Abhinaba
Puri, Geeta
Herremans, Dorien
contents We present Text2midi-InferAlign, a novel technique for improving symbolic music generation at inference time. Our method leverages text-to-audio alignment and music structural alignment rewards during inference to encourage the generated music to be consistent with the input caption. Specifically, we introduce two objectives scores: a text-audio consistency score that measures rhythmic alignment between the generated music and the original text caption, and a harmonic consistency score that penalizes generated music containing notes inconsistent with the key. By optimizing these alignment-based objectives during the generation process, our model produces symbolic music that is more closely tied to the input captions, thereby improving the overall quality and coherence of the generated compositions. Our approach can extend any existing autoregressive model without requiring further training or fine-tuning. We evaluate our work on top of Text2midi - an existing text-to-midi generation model, demonstrating significant improvements in both objective and subjective evaluation metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12669
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment
Roy, Abhinaba
Puri, Geeta
Herremans, Dorien
Sound
Artificial Intelligence
Multimedia
Audio and Speech Processing
68T07
I.2.1
We present Text2midi-InferAlign, a novel technique for improving symbolic music generation at inference time. Our method leverages text-to-audio alignment and music structural alignment rewards during inference to encourage the generated music to be consistent with the input caption. Specifically, we introduce two objectives scores: a text-audio consistency score that measures rhythmic alignment between the generated music and the original text caption, and a harmonic consistency score that penalizes generated music containing notes inconsistent with the key. By optimizing these alignment-based objectives during the generation process, our model produces symbolic music that is more closely tied to the input captions, thereby improving the overall quality and coherence of the generated compositions. Our approach can extend any existing autoregressive model without requiring further training or fine-tuning. We evaluate our work on top of Text2midi - an existing text-to-midi generation model, demonstrating significant improvements in both objective and subjective evaluation metrics.
title Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment
topic Sound
Artificial Intelligence
Multimedia
Audio and Speech Processing
68T07
I.2.1
url https://arxiv.org/abs/2505.12669