Integrating Text-to-Music Models with Language Models: Composing Long Structured Music Pieces

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autore principale: Atassi, Lilac
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909336914100224
author Atassi, Lilac
author_facet Atassi, Lilac
contents Recent music generation methods based on transformers have a context window of up to a minute. The music generated by these methods is largely unstructured beyond the context window. With a longer context window, learning long-scale structures from musical data is a prohibitively challenging problem. This paper proposes integrating a text-to-music model with a large language model to generate music with form. The papers discusses the solutions to the challenges of such integration. The experimental results show that the proposed method can generate 2.5-minute-long music that is highly structured, strongly organized, and cohesive.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00344
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Integrating Text-to-Music Models with Language Models: Composing Long Structured Music Pieces
Atassi, Lilac
Sound
Machine Learning
Audio and Speech Processing
Recent music generation methods based on transformers have a context window of up to a minute. The music generated by these methods is largely unstructured beyond the context window. With a longer context window, learning long-scale structures from musical data is a prohibitively challenging problem. This paper proposes integrating a text-to-music model with a large language model to generate music with form. The papers discusses the solutions to the challenges of such integration. The experimental results show that the proposed method can generate 2.5-minute-long music that is highly structured, strongly organized, and cohesive.
title Integrating Text-to-Music Models with Language Models: Composing Long Structured Music Pieces
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2410.00344