MART: Learning Hierarchical Music Audio Representations with Part-Whole Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Dong, Zhu, Jieming, Xun, Jiahao, Zhang, Shengyu, Zhao, Zhou, Deng, Liqun, Zhang, Wenqiao, Dong, Zhenhua, Jiang, Xin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911845587091456
author Yao, Dong
Zhu, Jieming
Xun, Jiahao
Zhang, Shengyu
Zhao, Zhou
Deng, Liqun
Zhang, Wenqiao
Dong, Zhenhua
Jiang, Xin
author_facet Yao, Dong
Zhu, Jieming
Xun, Jiahao
Zhang, Shengyu
Zhao, Zhou
Deng, Liqun
Zhang, Wenqiao
Dong, Zhenhua
Jiang, Xin
contents Recent research in self-supervised contrastive learning of music representations has demonstrated remarkable results across diverse downstream tasks. However, a prevailing trend in existing methods involves representing equally-sized music clips in either waveform or spectrogram formats, often overlooking the intrinsic part-whole hierarchies within music. In our quest to comprehend the bottom-up structure of music, we introduce MART, a hierarchical music representation learning approach that facilitates feature interactions among cropped music clips while considering their part-whole hierarchies. Specifically, we propose a hierarchical part-whole transformer to capture the structural relationships between music clips in a part-whole hierarchy. Furthermore, a hierarchical contrastive learning objective is crafted to align part-whole music representations at adjacent levels, progressively establishing a multi-hierarchy representation space. The effectiveness of our music representation learning from part-whole hierarchies has been empirically validated across multiple downstream tasks, including music classification and cover song identification.
format Preprint
id arxiv_https___arxiv_org_abs_2312_06197
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MART: Learning Hierarchical Music Audio Representations with Part-Whole Transformer
Yao, Dong
Zhu, Jieming
Xun, Jiahao
Zhang, Shengyu
Zhao, Zhou
Deng, Liqun
Zhang, Wenqiao
Dong, Zhenhua
Jiang, Xin
Sound
Multimedia
Audio and Speech Processing
Recent research in self-supervised contrastive learning of music representations has demonstrated remarkable results across diverse downstream tasks. However, a prevailing trend in existing methods involves representing equally-sized music clips in either waveform or spectrogram formats, often overlooking the intrinsic part-whole hierarchies within music. In our quest to comprehend the bottom-up structure of music, we introduce MART, a hierarchical music representation learning approach that facilitates feature interactions among cropped music clips while considering their part-whole hierarchies. Specifically, we propose a hierarchical part-whole transformer to capture the structural relationships between music clips in a part-whole hierarchy. Furthermore, a hierarchical contrastive learning objective is crafted to align part-whole music representations at adjacent levels, progressively establishing a multi-hierarchy representation space. The effectiveness of our music representation learning from part-whole hierarchies has been empirically validated across multiple downstream tasks, including music classification and cover song identification.
title MART: Learning Hierarchical Music Audio Representations with Part-Whole Transformer
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2312.06197