YuE: Scaling Open Foundation Models for Long-Form Music Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909787631910912 |
|---|---|
| author | Yuan, Ruibin Lin, Hanfeng Guo, Shuyue Zhang, Ge Pan, Jiahao Zang, Yongyi Liu, Haohe Liang, Yiming Ma, Wenye Du, Xingjian Du, Xinrun Ye, Zhen Zheng, Tianyu Jiang, Zhengxuan Ma, Yinghao Liu, Minghao Tian, Zeyue Zhou, Ziya Xue, Liumeng Qu, Xingwei Li, Yizhi Wu, Shangda Shen, Tianhao Ma, Ziyang Zhan, Jun Wang, Chunhui Wang, Yatian Chi, Xiaowei Zhang, Xinyue Yang, Zhenzhu Wang, Xiangzhou Liu, Shansong Mei, Lingrui Li, Peng Wang, Junjie Yu, Jianwei Pang, Guojian Li, Xu Wang, Zihao Zhou, Xiaohuan Yu, Lijun Benetos, Emmanouil Chen, Yong Lin, Chenghua Chen, Xie Xia, Gus Zhang, Zhaoxiang Zhang, Chao Chen, Wenhu Zhou, Xinyu Qiu, Xipeng Dannenberg, Roger Liu, Jiaheng Yang, Jian Huang, Wenhao Xue, Wei Tan, Xu Guo, Yike |
| author_facet | Yuan, Ruibin Lin, Hanfeng Guo, Shuyue Zhang, Ge Pan, Jiahao Zang, Yongyi Liu, Haohe Liang, Yiming Ma, Wenye Du, Xingjian Du, Xinrun Ye, Zhen Zheng, Tianyu Jiang, Zhengxuan Ma, Yinghao Liu, Minghao Tian, Zeyue Zhou, Ziya Xue, Liumeng Qu, Xingwei Li, Yizhi Wu, Shangda Shen, Tianhao Ma, Ziyang Zhan, Jun Wang, Chunhui Wang, Yatian Chi, Xiaowei Zhang, Xinyue Yang, Zhenzhu Wang, Xiangzhou Liu, Shansong Mei, Lingrui Li, Peng Wang, Junjie Yu, Jianwei Pang, Guojian Li, Xu Wang, Zihao Zhou, Xiaohuan Yu, Lijun Benetos, Emmanouil Chen, Yong Lin, Chenghua Chen, Xie Xia, Gus Zhang, Zhaoxiang Zhang, Chao Chen, Wenhu Zhou, Xinyu Qiu, Xipeng Dannenberg, Roger Liu, Jiaheng Yang, Jian Huang, Wenhao Xue, Wei Tan, Xu Guo, Yike |
| contents | We tackle the task of long-form music generation--particularly the challenging \textbf{lyrics-to-song} problem--by introducing YuE, a family of open foundation models based on the LLaMA2 architecture. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while maintaining lyrical alignment, coherent musical structure, and engaging vocal melodies with appropriate accompaniment. It achieves this through (1) track-decoupled next-token prediction to overcome dense mixture signals, (2) structural progressive conditioning for long-context lyrical alignment, and (3) a multitask, multiphase pre-training recipe to converge and generalize. In addition, we redesign the in-context learning technique for music generation, enabling versatile style transfer (e.g., converting Japanese city pop into an English rap while preserving the original accompaniment) and bidirectional generation. Through extensive evaluation, we demonstrate that YuE matches or even surpasses some of the proprietary systems in musicality and vocal agility. In addition, fine-tuning YuE enables additional controls and enhanced support for tail languages. Furthermore, beyond generation, we show that YuE's learned representations can perform well on music understanding tasks, where the results of YuE match or exceed state-of-the-art methods on the MARBLE benchmark. Keywords: lyrics2song, song generation, long-form, foundation model, music generation |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_08638 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | YuE: Scaling Open Foundation Models for Long-Form Music Generation Yuan, Ruibin Lin, Hanfeng Guo, Shuyue Zhang, Ge Pan, Jiahao Zang, Yongyi Liu, Haohe Liang, Yiming Ma, Wenye Du, Xingjian Du, Xinrun Ye, Zhen Zheng, Tianyu Jiang, Zhengxuan Ma, Yinghao Liu, Minghao Tian, Zeyue Zhou, Ziya Xue, Liumeng Qu, Xingwei Li, Yizhi Wu, Shangda Shen, Tianhao Ma, Ziyang Zhan, Jun Wang, Chunhui Wang, Yatian Chi, Xiaowei Zhang, Xinyue Yang, Zhenzhu Wang, Xiangzhou Liu, Shansong Mei, Lingrui Li, Peng Wang, Junjie Yu, Jianwei Pang, Guojian Li, Xu Wang, Zihao Zhou, Xiaohuan Yu, Lijun Benetos, Emmanouil Chen, Yong Lin, Chenghua Chen, Xie Xia, Gus Zhang, Zhaoxiang Zhang, Chao Chen, Wenhu Zhou, Xinyu Qiu, Xipeng Dannenberg, Roger Liu, Jiaheng Yang, Jian Huang, Wenhao Xue, Wei Tan, Xu Guo, Yike Audio and Speech Processing Artificial Intelligence Multimedia Sound We tackle the task of long-form music generation--particularly the challenging \textbf{lyrics-to-song} problem--by introducing YuE, a family of open foundation models based on the LLaMA2 architecture. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while maintaining lyrical alignment, coherent musical structure, and engaging vocal melodies with appropriate accompaniment. It achieves this through (1) track-decoupled next-token prediction to overcome dense mixture signals, (2) structural progressive conditioning for long-context lyrical alignment, and (3) a multitask, multiphase pre-training recipe to converge and generalize. In addition, we redesign the in-context learning technique for music generation, enabling versatile style transfer (e.g., converting Japanese city pop into an English rap while preserving the original accompaniment) and bidirectional generation. Through extensive evaluation, we demonstrate that YuE matches or even surpasses some of the proprietary systems in musicality and vocal agility. In addition, fine-tuning YuE enables additional controls and enhanced support for tail languages. Furthermore, beyond generation, we show that YuE's learned representations can perform well on music understanding tasks, where the results of YuE match or exceed state-of-the-art methods on the MARBLE benchmark. Keywords: lyrics2song, song generation, long-form, foundation model, music generation |
| title | YuE: Scaling Open Foundation Models for Long-Form Music Generation |
| topic | Audio and Speech Processing Artificial Intelligence Multimedia Sound |
| url | https://arxiv.org/abs/2503.08638 |