Text2midi: Generating Symbolic Music from Captions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bhandari, Keshav, Roy, Abhinaba, Wang, Kyra, Puri, Geeta, Colton, Simon, Herremans, Dorien
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911008097828864
author Bhandari, Keshav
Roy, Abhinaba
Wang, Kyra
Puri, Geeta
Colton, Simon
Herremans, Dorien
author_facet Bhandari, Keshav
Roy, Abhinaba
Wang, Kyra
Puri, Geeta
Colton, Simon
Herremans, Dorien
contents This paper introduces text2midi, an end-to-end model to generate MIDI files from textual descriptions. Leveraging the growing popularity of multimodal generative approaches, text2midi capitalizes on the extensive availability of textual data and the success of large language models (LLMs). Our end-to-end system harnesses the power of LLMs to generate symbolic music in the form of MIDI files. Specifically, we utilize a pretrained LLM encoder to process captions, which then condition an autoregressive transformer decoder to produce MIDI sequences that accurately reflect the provided descriptions. This intuitive and user-friendly method significantly streamlines the music creation process by allowing users to generate music pieces using text prompts. We conduct comprehensive empirical evaluations, incorporating both automated and human studies, that show our model generates MIDI files of high quality that are indeed controllable by text captions that may include music theory terms such as chords, keys, and tempo. We release the code and music samples on our demo page (https://github.com/AMAAI-Lab/Text2midi) for users to interact with text2midi.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16526
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Text2midi: Generating Symbolic Music from Captions
Bhandari, Keshav
Roy, Abhinaba
Wang, Kyra
Puri, Geeta
Colton, Simon
Herremans, Dorien
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
This paper introduces text2midi, an end-to-end model to generate MIDI files from textual descriptions. Leveraging the growing popularity of multimodal generative approaches, text2midi capitalizes on the extensive availability of textual data and the success of large language models (LLMs). Our end-to-end system harnesses the power of LLMs to generate symbolic music in the form of MIDI files. Specifically, we utilize a pretrained LLM encoder to process captions, which then condition an autoregressive transformer decoder to produce MIDI sequences that accurately reflect the provided descriptions. This intuitive and user-friendly method significantly streamlines the music creation process by allowing users to generate music pieces using text prompts. We conduct comprehensive empirical evaluations, incorporating both automated and human studies, that show our model generates MIDI files of high quality that are indeed controllable by text captions that may include music theory terms such as chords, keys, and tempo. We release the code and music samples on our demo page (https://github.com/AMAAI-Lab/Text2midi) for users to interact with text2midi.
title Text2midi: Generating Symbolic Music from Captions
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2412.16526