InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Chong, Ma, Yukun, Chen, Qian, Wang, Wen, Zhao, Shengkui, Pan, Zexu, Wang, Hao, Ni, Chongjia, Nguyen, Trung Hieu, Zhou, Kun, Jiang, Yidi, Tan, Chaohong, Gao, Zhifu, Du, Zhihao, Ma, Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909517784023040
author Zhang, Chong
Ma, Yukun
Chen, Qian
Wang, Wen
Zhao, Shengkui
Pan, Zexu
Wang, Hao
Ni, Chongjia
Nguyen, Trung Hieu
Zhou, Kun
Jiang, Yidi
Tan, Chaohong
Gao, Zhifu
Du, Zhihao
Ma, Bin
author_facet Zhang, Chong
Ma, Yukun
Chen, Qian
Wang, Wen
Zhao, Shengkui
Pan, Zexu
Wang, Hao
Ni, Chongjia
Nguyen, Trung Hieu
Zhou, Kun
Jiang, Yidi
Tan, Chaohong
Gao, Zhifu
Du, Zhihao
Ma, Bin
contents We introduce InspireMusic, a framework integrated super resolution and large language model for high-fidelity long-form music generation. A unified framework generates high-fidelity music, songs, and audio, which incorporates an autoregressive transformer with a super-resolution flow-matching model. This framework enables the controllable generation of high-fidelity long-form music at a higher sampling rate from both text and audio prompts. Our model differs from previous approaches, as we utilize an audio tokenizer with one codebook that contains richer semantic information, thereby reducing training costs and enhancing efficiency. This combination enables us to achieve high-quality audio generation with long-form coherence of up to $8$ minutes. Then, an autoregressive transformer model based on Qwen 2.5 predicts audio tokens. Next, we employ a super-resolution flow-matching model to generate high-sampling rate audio with fine-grained details learned from an acoustic codec model. Comprehensive experiments show that the InspireMusic-1.5B-Long model has a comparable performance to recent top-tier open-source systems, including MusicGen and Stable Audio 2.0, on subjective and objective evaluations. The code and pre-trained models are released at https://github.com/FunAudioLLM/InspireMusic.
format Preprint
id arxiv_https___arxiv_org_abs_2503_00084
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation
Zhang, Chong
Ma, Yukun
Chen, Qian
Wang, Wen
Zhao, Shengkui
Pan, Zexu
Wang, Hao
Ni, Chongjia
Nguyen, Trung Hieu
Zhou, Kun
Jiang, Yidi
Tan, Chaohong
Gao, Zhifu
Du, Zhihao
Ma, Bin
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
We introduce InspireMusic, a framework integrated super resolution and large language model for high-fidelity long-form music generation. A unified framework generates high-fidelity music, songs, and audio, which incorporates an autoregressive transformer with a super-resolution flow-matching model. This framework enables the controllable generation of high-fidelity long-form music at a higher sampling rate from both text and audio prompts. Our model differs from previous approaches, as we utilize an audio tokenizer with one codebook that contains richer semantic information, thereby reducing training costs and enhancing efficiency. This combination enables us to achieve high-quality audio generation with long-form coherence of up to $8$ minutes. Then, an autoregressive transformer model based on Qwen 2.5 predicts audio tokens. Next, we employ a super-resolution flow-matching model to generate high-sampling rate audio with fine-grained details learned from an acoustic codec model. Comprehensive experiments show that the InspireMusic-1.5B-Long model has a comparable performance to recent top-tier open-source systems, including MusicGen and Stable Audio 2.0, on subjective and objective evaluations. The code and pre-trained models are released at https://github.com/FunAudioLLM/InspireMusic.
title InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2503.00084