A Theory-Based Explainable Deep Learning Architecture for Music Emotion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fong, Hortense, Kumar, Vineet, Sudhir, K.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909286519537664
author Fong, Hortense
Kumar, Vineet
Sudhir, K.
author_facet Fong, Hortense
Kumar, Vineet
Sudhir, K.
contents This paper paper develops a theory-based, explainable deep learning convolutional neural network (CNN) classifier to predict the time-varying emotional response to music. We design novel CNN filters that leverage the frequency harmonics structure from acoustic physics known to impact the perception of musical features. Our theory-based model is more parsimonious, but provides comparable predictive performance to atheoretical deep learning models, while performing better than models using handcrafted features. Our model can be complemented with handcrafted features, but the performance improvement is marginal. Importantly, the harmonics-based structure placed on the CNN filters provides better explainability for how the model predicts emotional response (valence and arousal), because emotion is closely related to consonance--a perceptual feature defined by the alignment of harmonics. Finally, we illustrate the utility of our model with an application involving digital advertising. Motivated by YouTube mid-roll ads, we conduct a lab experiment in which we exogenously insert ads at different times within videos. We find that ads placed in emotionally similar contexts increase ad engagement (lower skip rates, higher brand recall rates). Ad insertion based on emotional similarity metrics predicted by our theory-based, explainable model produces comparable or better engagement relative to atheoretical models.
format Preprint
id arxiv_https___arxiv_org_abs_2408_07113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Theory-Based Explainable Deep Learning Architecture for Music Emotion
Fong, Hortense
Kumar, Vineet
Sudhir, K.
Sound
Artificial Intelligence
Human-Computer Interaction
Audio and Speech Processing
This paper paper develops a theory-based, explainable deep learning convolutional neural network (CNN) classifier to predict the time-varying emotional response to music. We design novel CNN filters that leverage the frequency harmonics structure from acoustic physics known to impact the perception of musical features. Our theory-based model is more parsimonious, but provides comparable predictive performance to atheoretical deep learning models, while performing better than models using handcrafted features. Our model can be complemented with handcrafted features, but the performance improvement is marginal. Importantly, the harmonics-based structure placed on the CNN filters provides better explainability for how the model predicts emotional response (valence and arousal), because emotion is closely related to consonance--a perceptual feature defined by the alignment of harmonics. Finally, we illustrate the utility of our model with an application involving digital advertising. Motivated by YouTube mid-roll ads, we conduct a lab experiment in which we exogenously insert ads at different times within videos. We find that ads placed in emotionally similar contexts increase ad engagement (lower skip rates, higher brand recall rates). Ad insertion based on emotional similarity metrics predicted by our theory-based, explainable model produces comparable or better engagement relative to atheoretical models.
title A Theory-Based Explainable Deep Learning Architecture for Music Emotion
topic Sound
Artificial Intelligence
Human-Computer Interaction
Audio and Speech Processing
url https://arxiv.org/abs/2408.07113