Music Foundation Model as Generic Booster for Music Downstream Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913860260200448 |
|---|---|
| author | Liao, WeiHsiang Takida, Yuhta Ikemiya, Yukara Zhong, Zhi Lai, Chieh-Hsin Fabbro, Giorgio Shimada, Kazuki Toyama, Keisuke Cheuk, Kinwai Martínez-Ramírez, Marco A. Takahashi, Shusuke Uhlich, Stefan Akama, Taketo Choi, Woosung Koyama, Yuichiro Mitsufuji, Yuki |
| author_facet | Liao, WeiHsiang Takida, Yuhta Ikemiya, Yukara Zhong, Zhi Lai, Chieh-Hsin Fabbro, Giorgio Shimada, Kazuki Toyama, Keisuke Cheuk, Kinwai Martínez-Ramírez, Marco A. Takahashi, Shusuke Uhlich, Stefan Akama, Taketo Choi, Woosung Koyama, Yuichiro Mitsufuji, Yuki |
| contents | We demonstrate the efficacy of using intermediate representations from a single foundation model to enhance various music downstream tasks. We introduce SoniDo, a music foundation model (MFM) designed to extract hierarchical features from target music samples. By leveraging hierarchical intermediate features, SoniDo constrains the information granularity, leading to improved performance across various downstream tasks including both understanding and generative tasks. We specifically evaluated this approach on representative tasks such as music tagging, music transcription, music source separation, and music mixing. Our results reveal that the features extracted from foundation models provide valuable enhancements in training downstream task models. This highlights the capability of using features extracted from music foundation models as a booster for downstream tasks. Our approach not only benefits existing task-specific models but also supports music downstream tasks constrained by data scarcity. This paves the way for more effective and accessible music processing solutions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_01135 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Music Foundation Model as Generic Booster for Music Downstream Tasks Liao, WeiHsiang Takida, Yuhta Ikemiya, Yukara Zhong, Zhi Lai, Chieh-Hsin Fabbro, Giorgio Shimada, Kazuki Toyama, Keisuke Cheuk, Kinwai Martínez-Ramírez, Marco A. Takahashi, Shusuke Uhlich, Stefan Akama, Taketo Choi, Woosung Koyama, Yuichiro Mitsufuji, Yuki Sound Information Retrieval Machine Learning Audio and Speech Processing We demonstrate the efficacy of using intermediate representations from a single foundation model to enhance various music downstream tasks. We introduce SoniDo, a music foundation model (MFM) designed to extract hierarchical features from target music samples. By leveraging hierarchical intermediate features, SoniDo constrains the information granularity, leading to improved performance across various downstream tasks including both understanding and generative tasks. We specifically evaluated this approach on representative tasks such as music tagging, music transcription, music source separation, and music mixing. Our results reveal that the features extracted from foundation models provide valuable enhancements in training downstream task models. This highlights the capability of using features extracted from music foundation models as a booster for downstream tasks. Our approach not only benefits existing task-specific models but also supports music downstream tasks constrained by data scarcity. This paves the way for more effective and accessible music processing solutions. |
| title | Music Foundation Model as Generic Booster for Music Downstream Tasks |
| topic | Sound Information Retrieval Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2411.01135 |