SAGE-Music: Low-Latency Symbolic Music Generation via Attribute-Specialized Key-Value Head Sharing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914093410025472 |
|---|---|
| author | Tan, Jiaye Luo, Haonan Song, Linfeng Chen, Shuaiqi Lyu, Yishan Zhong, Zian Wang, Roujia Jiang, Daniel Zhang, Haoran Bai, Jiaming Cheng, Haoran Liao, Q. Vera Dong, Hao-Wen |
| author_facet | Tan, Jiaye Luo, Haonan Song, Linfeng Chen, Shuaiqi Lyu, Yishan Zhong, Zian Wang, Roujia Jiang, Daniel Zhang, Haoran Bai, Jiaming Cheng, Haoran Liao, Q. Vera Dong, Hao-Wen |
| contents | Low-latency symbolic music generation is essential for real-time improvisation and human-AI co-creation. Existing transformer-based models, however, face a trade-off between inference speed and musical quality. Traditional acceleration techniques such as embedding pooling significantly degrade quality, while recently proposed Byte Pair Encoding (BPE) methods - though effective on single-track piano data - suffer large performance drops in multi-track settings, as revealed by our analysis. We propose Attribute-Specialized Key-Value Head Sharing (AS-KVHS), adapted to music's structured symbolic representation, achieving about 30% inference speedup with only a negligible (about 0.4%) quality drop in objective evaluations and slight improvements in subjective listening tests. Our main contributions are (1) the first systematic study of BPE's generalizability in multi-track symbolic music, and (2) the introduction of AS-KVHS for low-latency symbolic music generation. Beyond these, we also release SAGE-Music, an open-source benchmark that matches or surpasses state-of-the-art models in generation quality. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_00395 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SAGE-Music: Low-Latency Symbolic Music Generation via Attribute-Specialized Key-Value Head Sharing Tan, Jiaye Luo, Haonan Song, Linfeng Chen, Shuaiqi Lyu, Yishan Zhong, Zian Wang, Roujia Jiang, Daniel Zhang, Haoran Bai, Jiaming Cheng, Haoran Liao, Q. Vera Dong, Hao-Wen Sound Artificial Intelligence Machine Learning Audio and Speech Processing Low-latency symbolic music generation is essential for real-time improvisation and human-AI co-creation. Existing transformer-based models, however, face a trade-off between inference speed and musical quality. Traditional acceleration techniques such as embedding pooling significantly degrade quality, while recently proposed Byte Pair Encoding (BPE) methods - though effective on single-track piano data - suffer large performance drops in multi-track settings, as revealed by our analysis. We propose Attribute-Specialized Key-Value Head Sharing (AS-KVHS), adapted to music's structured symbolic representation, achieving about 30% inference speedup with only a negligible (about 0.4%) quality drop in objective evaluations and slight improvements in subjective listening tests. Our main contributions are (1) the first systematic study of BPE's generalizability in multi-track symbolic music, and (2) the introduction of AS-KVHS for low-latency symbolic music generation. Beyond these, we also release SAGE-Music, an open-source benchmark that matches or surpasses state-of-the-art models in generation quality. |
| title | SAGE-Music: Low-Latency Symbolic Music Generation via Attribute-Specialized Key-Value Head Sharing |
| topic | Sound Artificial Intelligence Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2510.00395 |