SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Chenyu, Wang, Shuai, Chen, Hangting, Yu, Jianwei, Tan, Wei, Gu, Rongzhi, Xu, Yaoxun, Zhou, Yizhi, Zhu, Haina, Li, Haizhou
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912206533165056
author Yang, Chenyu
Wang, Shuai
Chen, Hangting
Yu, Jianwei
Tan, Wei
Gu, Rongzhi
Xu, Yaoxun
Zhou, Yizhi
Zhu, Haina
Li, Haizhou
author_facet Yang, Chenyu
Wang, Shuai
Chen, Hangting
Yu, Jianwei
Tan, Wei
Gu, Rongzhi
Xu, Yaoxun
Zhou, Yizhi
Zhu, Haina
Li, Haizhou
contents The emergence of novel generative modeling paradigms, particularly audio language models, has significantly advanced the field of song generation. Although state-of-the-art models are capable of synthesizing both vocals and accompaniment tracks up to several minutes long concurrently, research about partial adjustments or editing of existing songs is still underexplored, which allows for more flexible and effective production. In this paper, we present SongEditor, the first song editing paradigm that introduces the editing capabilities into language-modeling song generation approaches, facilitating both segment-wise and track-wise modifications. SongEditor offers the flexibility to adjust lyrics, vocals, and accompaniments, as well as synthesizing songs from scratch. The core components of SongEditor include a music tokenizer, an autoregressive language model, and a diffusion generator, enabling generating an entire section, masked lyrics, or even separated vocals and background music. Extensive experiments demonstrate that the proposed SongEditor achieves exceptional performance in end-to-end song editing, as evidenced by both objective and subjective metrics. Audio samples are available in https://cypress-yang.github.io/SongEditor_demo/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13786
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
Yang, Chenyu
Wang, Shuai
Chen, Hangting
Yu, Jianwei
Tan, Wei
Gu, Rongzhi
Xu, Yaoxun
Zhou, Yizhi
Zhu, Haina
Li, Haizhou
Audio and Speech Processing
Sound
The emergence of novel generative modeling paradigms, particularly audio language models, has significantly advanced the field of song generation. Although state-of-the-art models are capable of synthesizing both vocals and accompaniment tracks up to several minutes long concurrently, research about partial adjustments or editing of existing songs is still underexplored, which allows for more flexible and effective production. In this paper, we present SongEditor, the first song editing paradigm that introduces the editing capabilities into language-modeling song generation approaches, facilitating both segment-wise and track-wise modifications. SongEditor offers the flexibility to adjust lyrics, vocals, and accompaniments, as well as synthesizing songs from scratch. The core components of SongEditor include a music tokenizer, an autoregressive language model, and a diffusion generator, enabling generating an entire section, masked lyrics, or even separated vocals and background music. Extensive experiments demonstrate that the proposed SongEditor achieves exceptional performance in end-to-end song editing, as evidenced by both objective and subjective metrics. Audio samples are available in https://cypress-yang.github.io/SongEditor_demo/.
title SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2412.13786