TEAdapter: Supply abundant guidance for controllable text-to-music generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zou, Jialing, Mei, Jiahao, Nan, Xudong, Li, Jinghua, Dong, Daoguo, He, Liang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909282662875136
author Zou, Jialing
Mei, Jiahao
Nan, Xudong
Li, Jinghua
Dong, Daoguo
He, Liang
author_facet Zou, Jialing
Mei, Jiahao
Nan, Xudong
Li, Jinghua
Dong, Daoguo
He, Liang
contents Although current text-guided music generation technology can cope with simple creative scenarios, achieving fine-grained control over individual text-modality conditions remains challenging as user demands become more intricate. Accordingly, we introduce the TEAcher Adapter (TEAdapter), a compact plugin designed to guide the generation process with diverse control information provided by users. In addition, we explore the controllable generation of extended music by leveraging TEAdapter control groups trained on data of distinct structural functionalities. In general, we consider controls over global, elemental, and structural levels. Experimental results demonstrate that the proposed TEAdapter enables multiple precise controls and ensures high-quality music generation. Our module is also lightweight and transferable to any diffusion model architecture. Available code and demos will be found soon at https://github.com/Ashley1101/TEAdapter.
format Preprint
id arxiv_https___arxiv_org_abs_2408_04865
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TEAdapter: Supply abundant guidance for controllable text-to-music generation
Zou, Jialing
Mei, Jiahao
Nan, Xudong
Li, Jinghua
Dong, Daoguo
He, Liang
Sound
Multimedia
Audio and Speech Processing
Although current text-guided music generation technology can cope with simple creative scenarios, achieving fine-grained control over individual text-modality conditions remains challenging as user demands become more intricate. Accordingly, we introduce the TEAcher Adapter (TEAdapter), a compact plugin designed to guide the generation process with diverse control information provided by users. In addition, we explore the controllable generation of extended music by leveraging TEAdapter control groups trained on data of distinct structural functionalities. In general, we consider controls over global, elemental, and structural levels. Experimental results demonstrate that the proposed TEAdapter enables multiple precise controls and ensures high-quality music generation. Our module is also lightweight and transferable to any diffusion model architecture. Available code and demos will be found soon at https://github.com/Ashley1101/TEAdapter.
title TEAdapter: Supply abundant guidance for controllable text-to-music generation
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2408.04865