MART: Learning Hierarchical Music Audio Representations with Part-Whole Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911845587091456 |
|---|---|
| author | Yao, Dong Zhu, Jieming Xun, Jiahao Zhang, Shengyu Zhao, Zhou Deng, Liqun Zhang, Wenqiao Dong, Zhenhua Jiang, Xin |
| author_facet | Yao, Dong Zhu, Jieming Xun, Jiahao Zhang, Shengyu Zhao, Zhou Deng, Liqun Zhang, Wenqiao Dong, Zhenhua Jiang, Xin |
| contents | Recent research in self-supervised contrastive learning of music representations has demonstrated remarkable results across diverse downstream tasks. However, a prevailing trend in existing methods involves representing equally-sized music clips in either waveform or spectrogram formats, often overlooking the intrinsic part-whole hierarchies within music. In our quest to comprehend the bottom-up structure of music, we introduce MART, a hierarchical music representation learning approach that facilitates feature interactions among cropped music clips while considering their part-whole hierarchies. Specifically, we propose a hierarchical part-whole transformer to capture the structural relationships between music clips in a part-whole hierarchy. Furthermore, a hierarchical contrastive learning objective is crafted to align part-whole music representations at adjacent levels, progressively establishing a multi-hierarchy representation space. The effectiveness of our music representation learning from part-whole hierarchies has been empirically validated across multiple downstream tasks, including music classification and cover song identification. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_06197 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | MART: Learning Hierarchical Music Audio Representations with Part-Whole Transformer Yao, Dong Zhu, Jieming Xun, Jiahao Zhang, Shengyu Zhao, Zhou Deng, Liqun Zhang, Wenqiao Dong, Zhenhua Jiang, Xin Sound Multimedia Audio and Speech Processing Recent research in self-supervised contrastive learning of music representations has demonstrated remarkable results across diverse downstream tasks. However, a prevailing trend in existing methods involves representing equally-sized music clips in either waveform or spectrogram formats, often overlooking the intrinsic part-whole hierarchies within music. In our quest to comprehend the bottom-up structure of music, we introduce MART, a hierarchical music representation learning approach that facilitates feature interactions among cropped music clips while considering their part-whole hierarchies. Specifically, we propose a hierarchical part-whole transformer to capture the structural relationships between music clips in a part-whole hierarchy. Furthermore, a hierarchical contrastive learning objective is crafted to align part-whole music representations at adjacent levels, progressively establishing a multi-hierarchy representation space. The effectiveness of our music representation learning from part-whole hierarchies has been empirically validated across multiple downstream tasks, including music classification and cover song identification. |
| title | MART: Learning Hierarchical Music Audio Representations with Part-Whole Transformer |
| topic | Sound Multimedia Audio and Speech Processing |
| url | https://arxiv.org/abs/2312.06197 |