MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909658261749760 |
|---|---|
| author | Tsai, Fang-Duo Wu, Shih-Lun Lee, Weijaw Yang, Sheng-Ping Chen, Bo-Rui Cheng, Hao-Chung Yang, Yi-Hsuan |
| author_facet | Tsai, Fang-Duo Wu, Shih-Lun Lee, Weijaw Yang, Sheng-Ping Chen, Bo-Rui Cheng, Hao-Chung Yang, Yi-Hsuan |
| contents | We propose MuseControlLite, a lightweight mechanism designed to fine-tune text-to-music generation models for precise conditioning using various time-varying musical attributes and reference audio signals. The key finding is that positional embeddings, which have been seldom used by text-to-music generation models in the conditioner for text conditions, are critical when the condition of interest is a function of time. Using melody control as an example, our experiments show that simply adding rotary positional embeddings to the decoupled cross-attention layers increases control accuracy from 56.6% to 61.1%, while requiring 6.75 times fewer trainable parameters than state-of-the-art fine-tuning mechanisms, using the same pre-trained diffusion Transformer model of Stable Audio Open. We evaluate various forms of musical attribute control, audio inpainting, and audio outpainting, demonstrating improved controllability over MusicGen-Large and Stable Audio Open ControlNet at a significantly lower fine-tuning cost, with only 85M trainble parameters. Source code, model checkpoints, and demo examples are available at: https://musecontrollite.github.io/web/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_18729 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners Tsai, Fang-Duo Wu, Shih-Lun Lee, Weijaw Yang, Sheng-Ping Chen, Bo-Rui Cheng, Hao-Chung Yang, Yi-Hsuan Sound Artificial Intelligence Audio and Speech Processing We propose MuseControlLite, a lightweight mechanism designed to fine-tune text-to-music generation models for precise conditioning using various time-varying musical attributes and reference audio signals. The key finding is that positional embeddings, which have been seldom used by text-to-music generation models in the conditioner for text conditions, are critical when the condition of interest is a function of time. Using melody control as an example, our experiments show that simply adding rotary positional embeddings to the decoupled cross-attention layers increases control accuracy from 56.6% to 61.1%, while requiring 6.75 times fewer trainable parameters than state-of-the-art fine-tuning mechanisms, using the same pre-trained diffusion Transformer model of Stable Audio Open. We evaluate various forms of musical attribute control, audio inpainting, and audio outpainting, demonstrating improved controllability over MusicGen-Large and Stable Audio Open ControlNet at a significantly lower fine-tuning cost, with only 85M trainble parameters. Source code, model checkpoints, and demo examples are available at: https://musecontrollite.github.io/web/. |
| title | MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners |
| topic | Sound Artificial Intelligence Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.18729 |