MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tsai, Fang-Duo, Wu, Shih-Lun, Lee, Weijaw, Yang, Sheng-Ping, Chen, Bo-Rui, Cheng, Hao-Chung, Yang, Yi-Hsuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909658261749760
author Tsai, Fang-Duo
Wu, Shih-Lun
Lee, Weijaw
Yang, Sheng-Ping
Chen, Bo-Rui
Cheng, Hao-Chung
Yang, Yi-Hsuan
author_facet Tsai, Fang-Duo
Wu, Shih-Lun
Lee, Weijaw
Yang, Sheng-Ping
Chen, Bo-Rui
Cheng, Hao-Chung
Yang, Yi-Hsuan
contents We propose MuseControlLite, a lightweight mechanism designed to fine-tune text-to-music generation models for precise conditioning using various time-varying musical attributes and reference audio signals. The key finding is that positional embeddings, which have been seldom used by text-to-music generation models in the conditioner for text conditions, are critical when the condition of interest is a function of time. Using melody control as an example, our experiments show that simply adding rotary positional embeddings to the decoupled cross-attention layers increases control accuracy from 56.6% to 61.1%, while requiring 6.75 times fewer trainable parameters than state-of-the-art fine-tuning mechanisms, using the same pre-trained diffusion Transformer model of Stable Audio Open. We evaluate various forms of musical attribute control, audio inpainting, and audio outpainting, demonstrating improved controllability over MusicGen-Large and Stable Audio Open ControlNet at a significantly lower fine-tuning cost, with only 85M trainble parameters. Source code, model checkpoints, and demo examples are available at: https://musecontrollite.github.io/web/.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18729
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners
Tsai, Fang-Duo
Wu, Shih-Lun
Lee, Weijaw
Yang, Sheng-Ping
Chen, Bo-Rui
Cheng, Hao-Chung
Yang, Yi-Hsuan
Sound
Artificial Intelligence
Audio and Speech Processing
We propose MuseControlLite, a lightweight mechanism designed to fine-tune text-to-music generation models for precise conditioning using various time-varying musical attributes and reference audio signals. The key finding is that positional embeddings, which have been seldom used by text-to-music generation models in the conditioner for text conditions, are critical when the condition of interest is a function of time. Using melody control as an example, our experiments show that simply adding rotary positional embeddings to the decoupled cross-attention layers increases control accuracy from 56.6% to 61.1%, while requiring 6.75 times fewer trainable parameters than state-of-the-art fine-tuning mechanisms, using the same pre-trained diffusion Transformer model of Stable Audio Open. We evaluate various forms of musical attribute control, audio inpainting, and audio outpainting, demonstrating improved controllability over MusicGen-Large and Stable Audio Open ControlNet at a significantly lower fine-tuning cost, with only 85M trainble parameters. Source code, model checkpoints, and demo examples are available at: https://musecontrollite.github.io/web/.
title MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2506.18729