SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Chenyu, Wang, Shuai, Chen, Hangting, Tan, Wei, Yu, Jianwei, Li, Haizhou
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908605235593216
author Yang, Chenyu
Wang, Shuai
Chen, Hangting
Tan, Wei
Yu, Jianwei
Li, Haizhou
author_facet Yang, Chenyu
Wang, Shuai
Chen, Hangting
Tan, Wei
Yu, Jianwei
Li, Haizhou
contents Generating music with coherent structure, harmonious instrumental and vocal elements remains a significant challenge in song generation. Existing language models and diffusion-based methods often struggle to balance global coherence with local fidelity, resulting in outputs that lack musicality or suffer from incoherent progression and mismatched lyrics. This paper introduces $\textbf{SongBloom}$, a novel framework for full-length song generation that leverages an interleaved paradigm of autoregressive sketching and diffusion-based refinement. SongBloom employs an autoregressive diffusion model that combines the high fidelity of diffusion models with the scalability of language models. Specifically, it gradually extends a musical sketch from short to long and refines the details from coarse to fine-grained. The interleaved generation paradigm effectively integrates prior semantic and acoustic context to guide the generation process. Experimental results demonstrate that SongBloom outperforms existing methods across both subjective and objective metrics and achieves performance comparable to the state-of-the-art commercial music generation platforms. Audio samples are available on our demo page: https://cypress-yang.github.io/SongBloom_demo. The code and model weights have been released on https://github.com/Cypress-Yang/SongBloom .
format Preprint
id arxiv_https___arxiv_org_abs_2506_07634
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
Yang, Chenyu
Wang, Shuai
Chen, Hangting
Tan, Wei
Yu, Jianwei
Li, Haizhou
Audio and Speech Processing
Multimedia
Generating music with coherent structure, harmonious instrumental and vocal elements remains a significant challenge in song generation. Existing language models and diffusion-based methods often struggle to balance global coherence with local fidelity, resulting in outputs that lack musicality or suffer from incoherent progression and mismatched lyrics. This paper introduces $\textbf{SongBloom}$, a novel framework for full-length song generation that leverages an interleaved paradigm of autoregressive sketching and diffusion-based refinement. SongBloom employs an autoregressive diffusion model that combines the high fidelity of diffusion models with the scalability of language models. Specifically, it gradually extends a musical sketch from short to long and refines the details from coarse to fine-grained. The interleaved generation paradigm effectively integrates prior semantic and acoustic context to guide the generation process. Experimental results demonstrate that SongBloom outperforms existing methods across both subjective and objective metrics and achieves performance comparable to the state-of-the-art commercial music generation platforms. Audio samples are available on our demo page: https://cypress-yang.github.io/SongBloom_demo. The code and model weights have been released on https://github.com/Cypress-Yang/SongBloom .
title SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
topic Audio and Speech Processing
Multimedia
url https://arxiv.org/abs/2506.07634