MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Shih-Lun, Kim, Yoon, Huang, Cheng-Zhi Anna
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909889558740992
author Wu, Shih-Lun
Kim, Yoon
Huang, Cheng-Zhi Anna
author_facet Wu, Shih-Lun
Kim, Yoon
Huang, Cheng-Zhi Anna
contents We present MIDI-LLM, an LLM for generating multitrack MIDI music from free-form text prompts. Our approach expands a text LLM's vocabulary to include MIDI tokens, and uses a two-stage training recipe to endow text-to-MIDI abilities. By preserving the original LLM's parameter structure, we can directly leverage the vLLM library for accelerated inference. Experiments show that MIDI-LLM achieves higher quality, better text control, and faster inference compared to the recent Text2midi model. Live demo at https://midi-llm-demo.vercel.app.
format Preprint
id arxiv_https___arxiv_org_abs_2511_03942
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
Wu, Shih-Lun
Kim, Yoon
Huang, Cheng-Zhi Anna
Sound
Computation and Language
Multimedia
We present MIDI-LLM, an LLM for generating multitrack MIDI music from free-form text prompts. Our approach expands a text LLM's vocabulary to include MIDI tokens, and uses a two-stage training recipe to endow text-to-MIDI abilities. By preserving the original LLM's parameter structure, we can directly leverage the vLLM library for accelerated inference. Experiments show that MIDI-LLM achieves higher quality, better text control, and faster inference compared to the recent Text2midi model. Live demo at https://midi-llm-demo.vercel.app.
title MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
topic Sound
Computation and Language
Multimedia
url https://arxiv.org/abs/2511.03942