A Survey on Multimodal Music Emotion Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liyanarachchi, Rashini, Joshi, Aditya, Meijering, Erik
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910919417659392
author Liyanarachchi, Rashini
Joshi, Aditya
Meijering, Erik
author_facet Liyanarachchi, Rashini
Joshi, Aditya
Meijering, Erik
contents Multimodal music emotion recognition (MMER) is an emerging discipline in music information retrieval that has experienced a surge in interest in recent years. This survey provides a comprehensive overview of the current state-of-the-art in MMER. Discussing the different approaches and techniques used in this field, the paper introduces a four-stage MMER framework, including multimodal data selection, feature extraction, feature processing, and final emotion prediction. The survey further reveals significant advancements in deep learning methods and the increasing importance of feature fusion techniques. Despite these advancements, challenges such as the need for large annotated datasets, datasets with more modalities, and real-time processing capabilities remain. This paper also contributes to the field by identifying critical gaps in current research and suggesting potential directions for future research. The gaps underscore the importance of developing robust, scalable, a interpretable models for MMER, with implications for applications in music recommendation systems, therapeutic tools, and entertainment.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18799
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey on Multimodal Music Emotion Recognition
Liyanarachchi, Rashini
Joshi, Aditya
Meijering, Erik
Multimedia
Sound
Audio and Speech Processing
Multimodal music emotion recognition (MMER) is an emerging discipline in music information retrieval that has experienced a surge in interest in recent years. This survey provides a comprehensive overview of the current state-of-the-art in MMER. Discussing the different approaches and techniques used in this field, the paper introduces a four-stage MMER framework, including multimodal data selection, feature extraction, feature processing, and final emotion prediction. The survey further reveals significant advancements in deep learning methods and the increasing importance of feature fusion techniques. Despite these advancements, challenges such as the need for large annotated datasets, datasets with more modalities, and real-time processing capabilities remain. This paper also contributes to the field by identifying critical gaps in current research and suggesting potential directions for future research. The gaps underscore the importance of developing robust, scalable, a interpretable models for MMER, with implications for applications in music recommendation systems, therapeutic tools, and entertainment.
title A Survey on Multimodal Music Emotion Recognition
topic Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2504.18799