_version_ 1866909787631910912
author Yuan, Ruibin
Lin, Hanfeng
Guo, Shuyue
Zhang, Ge
Pan, Jiahao
Zang, Yongyi
Liu, Haohe
Liang, Yiming
Ma, Wenye
Du, Xingjian
Du, Xinrun
Ye, Zhen
Zheng, Tianyu
Jiang, Zhengxuan
Ma, Yinghao
Liu, Minghao
Tian, Zeyue
Zhou, Ziya
Xue, Liumeng
Qu, Xingwei
Li, Yizhi
Wu, Shangda
Shen, Tianhao
Ma, Ziyang
Zhan, Jun
Wang, Chunhui
Wang, Yatian
Chi, Xiaowei
Zhang, Xinyue
Yang, Zhenzhu
Wang, Xiangzhou
Liu, Shansong
Mei, Lingrui
Li, Peng
Wang, Junjie
Yu, Jianwei
Pang, Guojian
Li, Xu
Wang, Zihao
Zhou, Xiaohuan
Yu, Lijun
Benetos, Emmanouil
Chen, Yong
Lin, Chenghua
Chen, Xie
Xia, Gus
Zhang, Zhaoxiang
Zhang, Chao
Chen, Wenhu
Zhou, Xinyu
Qiu, Xipeng
Dannenberg, Roger
Liu, Jiaheng
Yang, Jian
Huang, Wenhao
Xue, Wei
Tan, Xu
Guo, Yike
author_facet Yuan, Ruibin
Lin, Hanfeng
Guo, Shuyue
Zhang, Ge
Pan, Jiahao
Zang, Yongyi
Liu, Haohe
Liang, Yiming
Ma, Wenye
Du, Xingjian
Du, Xinrun
Ye, Zhen
Zheng, Tianyu
Jiang, Zhengxuan
Ma, Yinghao
Liu, Minghao
Tian, Zeyue
Zhou, Ziya
Xue, Liumeng
Qu, Xingwei
Li, Yizhi
Wu, Shangda
Shen, Tianhao
Ma, Ziyang
Zhan, Jun
Wang, Chunhui
Wang, Yatian
Chi, Xiaowei
Zhang, Xinyue
Yang, Zhenzhu
Wang, Xiangzhou
Liu, Shansong
Mei, Lingrui
Li, Peng
Wang, Junjie
Yu, Jianwei
Pang, Guojian
Li, Xu
Wang, Zihao
Zhou, Xiaohuan
Yu, Lijun
Benetos, Emmanouil
Chen, Yong
Lin, Chenghua
Chen, Xie
Xia, Gus
Zhang, Zhaoxiang
Zhang, Chao
Chen, Wenhu
Zhou, Xinyu
Qiu, Xipeng
Dannenberg, Roger
Liu, Jiaheng
Yang, Jian
Huang, Wenhao
Xue, Wei
Tan, Xu
Guo, Yike
contents We tackle the task of long-form music generation--particularly the challenging \textbf{lyrics-to-song} problem--by introducing YuE, a family of open foundation models based on the LLaMA2 architecture. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while maintaining lyrical alignment, coherent musical structure, and engaging vocal melodies with appropriate accompaniment. It achieves this through (1) track-decoupled next-token prediction to overcome dense mixture signals, (2) structural progressive conditioning for long-context lyrical alignment, and (3) a multitask, multiphase pre-training recipe to converge and generalize. In addition, we redesign the in-context learning technique for music generation, enabling versatile style transfer (e.g., converting Japanese city pop into an English rap while preserving the original accompaniment) and bidirectional generation. Through extensive evaluation, we demonstrate that YuE matches or even surpasses some of the proprietary systems in musicality and vocal agility. In addition, fine-tuning YuE enables additional controls and enhanced support for tail languages. Furthermore, beyond generation, we show that YuE's learned representations can perform well on music understanding tasks, where the results of YuE match or exceed state-of-the-art methods on the MARBLE benchmark. Keywords: lyrics2song, song generation, long-form, foundation model, music generation
format Preprint
id arxiv_https___arxiv_org_abs_2503_08638
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle YuE: Scaling Open Foundation Models for Long-Form Music Generation
Yuan, Ruibin
Lin, Hanfeng
Guo, Shuyue
Zhang, Ge
Pan, Jiahao
Zang, Yongyi
Liu, Haohe
Liang, Yiming
Ma, Wenye
Du, Xingjian
Du, Xinrun
Ye, Zhen
Zheng, Tianyu
Jiang, Zhengxuan
Ma, Yinghao
Liu, Minghao
Tian, Zeyue
Zhou, Ziya
Xue, Liumeng
Qu, Xingwei
Li, Yizhi
Wu, Shangda
Shen, Tianhao
Ma, Ziyang
Zhan, Jun
Wang, Chunhui
Wang, Yatian
Chi, Xiaowei
Zhang, Xinyue
Yang, Zhenzhu
Wang, Xiangzhou
Liu, Shansong
Mei, Lingrui
Li, Peng
Wang, Junjie
Yu, Jianwei
Pang, Guojian
Li, Xu
Wang, Zihao
Zhou, Xiaohuan
Yu, Lijun
Benetos, Emmanouil
Chen, Yong
Lin, Chenghua
Chen, Xie
Xia, Gus
Zhang, Zhaoxiang
Zhang, Chao
Chen, Wenhu
Zhou, Xinyu
Qiu, Xipeng
Dannenberg, Roger
Liu, Jiaheng
Yang, Jian
Huang, Wenhao
Xue, Wei
Tan, Xu
Guo, Yike
Audio and Speech Processing
Artificial Intelligence
Multimedia
Sound
We tackle the task of long-form music generation--particularly the challenging \textbf{lyrics-to-song} problem--by introducing YuE, a family of open foundation models based on the LLaMA2 architecture. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while maintaining lyrical alignment, coherent musical structure, and engaging vocal melodies with appropriate accompaniment. It achieves this through (1) track-decoupled next-token prediction to overcome dense mixture signals, (2) structural progressive conditioning for long-context lyrical alignment, and (3) a multitask, multiphase pre-training recipe to converge and generalize. In addition, we redesign the in-context learning technique for music generation, enabling versatile style transfer (e.g., converting Japanese city pop into an English rap while preserving the original accompaniment) and bidirectional generation. Through extensive evaluation, we demonstrate that YuE matches or even surpasses some of the proprietary systems in musicality and vocal agility. In addition, fine-tuning YuE enables additional controls and enhanced support for tail languages. Furthermore, beyond generation, we show that YuE's learned representations can perform well on music understanding tasks, where the results of YuE match or exceed state-of-the-art methods on the MARBLE benchmark. Keywords: lyrics2song, song generation, long-form, foundation model, music generation
title YuE: Scaling Open Foundation Models for Long-Form Music Generation
topic Audio and Speech Processing
Artificial Intelligence
Multimedia
Sound
url https://arxiv.org/abs/2503.08638