Semantic-Aware Interpretable Multimodal Music Auto-Tagging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Patakis, Andreas, Lyberatos, Vassilis, Kantarelis, Spyridon, Dervakos, Edmund, Stamou, Giorgos
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914605851213824
author Patakis, Andreas
Lyberatos, Vassilis
Kantarelis, Spyridon
Dervakos, Edmund
Stamou, Giorgos
author_facet Patakis, Andreas
Lyberatos, Vassilis
Kantarelis, Spyridon
Dervakos, Edmund
Stamou, Giorgos
contents Music auto-tagging is essential for organizing and discovering music in extensive digital libraries. While foundation models achieve exceptional performance in this domain, their outputs often lack interpretability, limiting trust and usability for researchers and end-users alike. In this work, we present an interpretable framework for music auto-tagging that leverages groups of musically meaningful multimodal features, derived from signal processing, deep learning, ontology engineering, and natural language processing. To enhance interpretability, we cluster features semantically and employ an expectation maximization algorithm, assigning distinct weights to each group based on its contribution to the tagging process. Our method achieves competitive tagging performance while offering a deeper understanding of the decision-making process, paving the way for more transparent and user-centric music tagging systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17233
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semantic-Aware Interpretable Multimodal Music Auto-Tagging
Patakis, Andreas
Lyberatos, Vassilis
Kantarelis, Spyridon
Dervakos, Edmund
Stamou, Giorgos
Machine Learning
Sound
Audio and Speech Processing
Music auto-tagging is essential for organizing and discovering music in extensive digital libraries. While foundation models achieve exceptional performance in this domain, their outputs often lack interpretability, limiting trust and usability for researchers and end-users alike. In this work, we present an interpretable framework for music auto-tagging that leverages groups of musically meaningful multimodal features, derived from signal processing, deep learning, ontology engineering, and natural language processing. To enhance interpretability, we cluster features semantically and employ an expectation maximization algorithm, assigning distinct weights to each group based on its contribution to the tagging process. Our method achieves competitive tagging performance while offering a deeper understanding of the decision-making process, paving the way for more transparent and user-centric music tagging systems.
title Semantic-Aware Interpretable Multimodal Music Auto-Tagging
topic Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.17233